How to Build an ETL App for HubDB Data in Python with CData Connect AI
The rich ecosystem of Python modules lets you get to work quickly and integrate your systems more effectively. With the CData Connect AI Python SDK and the petl framework, you can build HubDB-connected applications and pipelines for extracting, transforming, and loading HubDB data. This article shows how to connect to Connect AI and use petl to extract, transform, and load HubDB data.
The Connect AI Python SDK (cdata-connect-ai) is a DB-API 2.0 (PEP 249) compliant client, so petl can read directly from the SDK connection with etl.fromdb. There is no driver to install per source: connect with a Personal Access Token and build your pipeline.
Connect to HubDB in Connect AI
CData Connect AI uses a straightforward, point-and-click interface to connect to data sources.
- Log into Connect AI, click Sources, and then click Add Connection
- Select "HubDB" from the Add Connection panel
-
Enter the necessary authentication properties to connect to HubDB.
There are two authentication methods available for connecting to HubDB data source: OAuth Authentication with a public HubSpot application and authentication with a Private application token.
Using a Custom OAuth App
AuthScheme must be set to "OAuth" in all OAuth flows. Be sure to review the Help documentation for the required connection properties for you specific authentication needs (desktop applications, web applications, and headless machines).
Follow the steps below to register an application and obtain the OAuth client credentials:
- Log into your HubSpot app developer account.
- Note that it must be an app developer account. Standard HubSpot accounts cannot create public apps.
- On the developer account home page, click the Apps tab.
- Click Create app.
- On the App info tab, enter and optionally modify values that are displayed to users when they connect. These values include the public application name, application logo, and a description of the application.
- On the Auth tab, supply a callback URL in the "Redirect URLs" box.
- If you're creating a desktop application, set this to a locally accessible URL like http://localhost:33333.
- If you are creating a Web application, set this to a trusted URL where you want users to be redirected to when they authorize your application.
- Click Create App. HubSpot then generates the application, along with its associated credentials.
- On the Auth tab, note the Client ID and Client secret. You will use these later to configure the driver.
Under Scopes, select any scopes you need for your application's intended functionality.
A minimum of the following scopes is required to access tables:
- hubdb
- oauth
- crm.objects.owners.read
- Click Save changes.
- Install the application into a production portal with access to the features that are required by the integration.
- Under "Install URL (OAuth)", click Copy full URL to copy the installation URL for your application.
- Navigate to the copied link in your browser. Select a standard account in which to install the application.
- Click Connect app. You can close the resulting tab.
Using a Private App
To connect using a HubSpot private application token, set the AuthScheme property to "PrivateApp."
You can generate a private application token by following the steps below:
- In your HubDB account, click the settings icon (the gear) in the main navigation bar.
- In the left sidebar menu, navigate to Integrations > Private Apps.
- Click Create private app.
- On the Basic Info tab, configure the details of your application (name, logo, and description).
- On the Scopes tab, select Read or Write for each scope you want your private application to be able to access.
- A minimum of hubdb and crm.objects.owners.read is required to access tables.
- After you are done configuring your application, click Create app in the top right.
- Review the info about your application's access token, click Continue creating, and then Show token.
- Click Copy to copy the private application token.
To connect, set PrivateAppToken to the private application token you retrieved.
- Log into your HubSpot app developer account.
- Click Save & Test
- Navigate to the Permissions tab and update the user-based permissions.

Generate a Personal Access Token (PAT)
The Python SDK authenticates to Connect AI with your account email and a Personal Access Token (PAT). It is best practice to create a separate PAT for each application to maintain granularity of access.
- Click the Gear icon () at the top right of the Connect AI app to open the Settings page.
- On the Settings page, go to the Access Tokens section and click Create PAT.
- Give the PAT a name and click Create.

- The PAT is only visible at creation, so copy it and store it securely.
Install Required Modules
Install the SDK and the petl framework using the pip utility:
pip install cdata-connect-ai pip install petl
Build an ETL App for HubDB Data in Python
Once the required modules are installed, you are ready to build the ETL app. Code snippets follow, but the full source code is available at the end of the article.
First, import the modules and connect to Connect AI with your account email and PAT:
import petl as etl
import cdata_connect_ai
conn = cdata_connect_ai.connect(
username="[email protected]",
password="<your_pat>",
)
Create a SQL Statement to Query HubDB
Use SQL to create a statement for querying HubDB. In this article, we read data from the NorthwindProducts entity. Identifiers are three-part: <Connection>.<Schema>.<Table>, where the connection name defaults to the source name (for example, HubDB1).
sql = (
"SELECT PartitionKey, Name "
"FROM [HubDB1].[HubDB].[NorthwindProducts] "
"WHERE Id = '1'"
)
Extract, Transform, and Load the HubDB Data
With a connection and query in hand, use petl to extract, transform, and load the HubDB data. In this example, we extract HubDB data, sort the data by the Name column, and load the data into a CSV file.
table1 = etl.fromdb(conn, sql) table2 = etl.sort(table1, 'Name') etl.tocsv(table2, 'northwindproducts_data.csv')
Load New Rows Back into HubDB
When HubDB supports writes, load rows back with a batch INSERT. The SDK's executemany takes @name placeholders and a list of parameter dictionaries, one per row.
cur = conn.cursor()
cur.executemany(
"INSERT INTO [HubDB1].[HubDB].[NorthwindProducts] (PartitionKey, Name) "
"VALUES (@val1, @val2)",
[
{"@val1": "New value 1", "@val2": "New value 1"},
{"@val1": "New value 2", "@val2": "New value 2"},
],
)
print(f"Rows inserted: {cur.rowcount}")
conn.close()
Note: Even for writable sources, a read-only PAT or connection permission will reject write operations.
With the CData Connect AI Python SDK, you can work with HubDB data just like you would with any database, including direct access to data in ETL packages like petl.
More Information and Free Trial
Now you can pipe live HubDB data through petl using the CData Connect AI Python SDK. For more information on connecting to HubDB (and hundreds of other data sources), visit the Connect AI page. Sign up for a free trial and start building data pipelines for live HubDB data in Python.
Full Source Code
import petl as etl
import cdata_connect_ai
conn = cdata_connect_ai.connect(
username="[email protected]",
password="<your_pat>",
)
sql = (
"SELECT PartitionKey, Name "
"FROM [HubDB1].[HubDB].[NorthwindProducts] "
"WHERE Id = '1'"
)
table1 = etl.fromdb(conn, sql)
table2 = etl.sort(table1, 'Name')
etl.tocsv(table2, 'northwindproducts_data.csv')
cur = conn.cursor()
cur.executemany(
"INSERT INTO [HubDB1].[HubDB].[NorthwindProducts] (PartitionKey, Name) "
"VALUES (@val1, @val2)",
[
{"@val1": "New value 1", "@val2": "New value 1"},
{"@val1": "New value 2", "@val2": "New value 2"},
],
)
print(f"Rows inserted: {cur.rowcount}")
conn.close()