How to Build an ETL App for SAP SuccessFactors Data in Python with CData Connect AI
The rich ecosystem of Python modules lets you get to work quickly and integrate your systems more effectively. With the CData Connect AI Python SDK and the petl framework, you can build SAP SuccessFactors-connected applications and pipelines for extracting, transforming, and loading SAP SuccessFactors data. This article shows how to connect to Connect AI and use petl to extract, transform, and load SAP SuccessFactors data.
The Connect AI Python SDK (cdata-connect-ai) is a DB-API 2.0 (PEP 249) compliant client, so petl can read directly from the SDK connection with etl.fromdb. There is no driver to install per source: connect with a Personal Access Token and build your pipeline.
Connect to SAP SuccessFactors in Connect AI
CData Connect AI uses a straightforward, point-and-click interface to connect to data sources.
- Log into Connect AI, click Sources, and then click Add Connection
- Select "SAP SuccessFactors" from the Add Connection panel
-
Enter the necessary authentication properties to connect to SAP SuccessFactors.
You can authenticate to SAP Success Factors using Basic authentication or OAuth with SAML assertion.
Basic Authentication
You must provide values for the following properties to successfully authenticate to SAP Success Factors. Note that the provider will reuse the session opened by SAP Success Factors using cookies. Which means that your credentials will be used only on the first request to open the session. After that, cookies returned from SAP Success Factors will be used for authentication.
- Url: set this to the URL of the server hosting Success Factors. Some of the servers are listed in the SAP support documentation (external link).
- User: set this to the username of your account.
- Password: set this to the password of your account.
- CompanyId: set this to the unique identifier of your company.
OAuth Authentication
You must provide values for the following properties, which will be used to get the access token.
- Url: set this to the URL of the server hosting Success Factors. Some of the servers are listed in the SAP support documentation (external link).
- User: set this to the username of your account.
- CompanyId: set this to the unique identifier of your company.
- OAuthClientId: set this to the API Key that was generated in API Center.
- OAuthClientSecret: the X.509 private key used to sign SAML assertion. The private key can be found in the certificate you downloaded in Registering your OAuth Client Application.
- InitiateOAuth: set this to GETANDREFRESH.
- Click Save & Test
- Navigate to the Permissions tab and update the user-based permissions.

Generate a Personal Access Token (PAT)
The Python SDK authenticates to Connect AI with your account email and a Personal Access Token (PAT). It is best practice to create a separate PAT for each application to maintain granularity of access.
- Click the Gear icon () at the top right of the Connect AI app to open the Settings page.
- On the Settings page, go to the Access Tokens section and click Create PAT.
- Give the PAT a name and click Create.

- The PAT is only visible at creation, so copy it and store it securely.
Install Required Modules
Install the SDK and the petl framework using the pip utility:
pip install cdata-connect-ai pip install petl
Build an ETL App for SAP SuccessFactors Data in Python
Once the required modules are installed, you are ready to build the ETL app. Code snippets follow, but the full source code is available at the end of the article.
First, import the modules and connect to Connect AI with your account email and PAT:
import petl as etl
import cdata_connect_ai
conn = cdata_connect_ai.connect(
username="[email protected]",
password="<your_pat>",
)
Create a SQL Statement to Query SAP SuccessFactors
Use SQL to create a statement for querying SAP SuccessFactors. In this article, we read data from the ExtAddressInfo entity. Identifiers are three-part: <Connection>.<Schema>.<Table>, where the connection name defaults to the source name (for example, SAPSuccessFactors1).
sql = (
"SELECT address1, zipCode "
"FROM [SAPSuccessFactors1].[SAPSuccessFactors].[ExtAddressInfo] "
"WHERE city = 'Springfield'"
)
Extract, Transform, and Load the SAP SuccessFactors Data
With a connection and query in hand, use petl to extract, transform, and load the SAP SuccessFactors data. In this example, we extract SAP SuccessFactors data, sort the data by the zipCode column, and load the data into a CSV file.
table1 = etl.fromdb(conn, sql) table2 = etl.sort(table1, 'zipCode') etl.tocsv(table2, 'extaddressinfo_data.csv')
Load New Rows Back into SAP SuccessFactors
When SAP SuccessFactors supports writes, load rows back with a batch INSERT. The SDK's executemany takes @name placeholders and a list of parameter dictionaries, one per row.
cur = conn.cursor()
cur.executemany(
"INSERT INTO [SAPSuccessFactors1].[SAPSuccessFactors].[ExtAddressInfo] (address1, zipCode) "
"VALUES (@val1, @val2)",
[
{"@val1": "New value 1", "@val2": "New value 1"},
{"@val1": "New value 2", "@val2": "New value 2"},
],
)
print(f"Rows inserted: {cur.rowcount}")
conn.close()
Note: Even for writable sources, a read-only PAT or connection permission will reject write operations.
With the CData Connect AI Python SDK, you can work with SAP SuccessFactors data just like you would with any database, including direct access to data in ETL packages like petl.
More Information and Free Trial
Now you can pipe live SAP SuccessFactors data through petl using the CData Connect AI Python SDK. For more information on connecting to SAP SuccessFactors (and hundreds of other data sources), visit the Connect AI page. Sign up for a free trial and start building data pipelines for live SAP SuccessFactors data in Python.
Full Source Code
import petl as etl
import cdata_connect_ai
conn = cdata_connect_ai.connect(
username="[email protected]",
password="<your_pat>",
)
sql = (
"SELECT address1, zipCode "
"FROM [SAPSuccessFactors1].[SAPSuccessFactors].[ExtAddressInfo] "
"WHERE city = 'Springfield'"
)
table1 = etl.fromdb(conn, sql)
table2 = etl.sort(table1, 'zipCode')
etl.tocsv(table2, 'extaddressinfo_data.csv')
cur = conn.cursor()
cur.executemany(
"INSERT INTO [SAPSuccessFactors1].[SAPSuccessFactors].[ExtAddressInfo] (address1, zipCode) "
"VALUES (@val1, @val2)",
[
{"@val1": "New value 1", "@val2": "New value 1"},
{"@val1": "New value 2", "@val2": "New value 2"},
],
)
print(f"Rows inserted: {cur.rowcount}")
conn.close()