Getting Started with the CData Connect AI Python SDK for Google Cloud Storage
The CData Connect AI Python SDK (cdata-connect-ai) is a DB-API 2.0 (PEP 249) compliant client that lets you fetch and act on live Google Cloud Storage data with standard Python database code. Because Connect AI provides the connectivity: you install one package, authenticate with a Personal Access Token, and query Google Cloud Storage (and every other source connected in Connect AI) using the same familiar connect() / cursor() / fetchall() pattern you already know from libraries like sqlite3 and psycopg2.
This guide walks through connecting Google Cloud Storage in Connect AI, generating a Personal Access Token, installing the SDK, and reading (and, where supported, writing) live Google Cloud Storage data.
Prerequisites
- An account in CData Connect AI
- Python 3.8 or higher
- An active Google Cloud Storage account with valid credentials
Connect to Google Cloud Storage in Connect AI
CData Connect AI uses a straightforward, point-and-click interface to connect to data sources.
- Log into Connect AI, click Sources, and then click Add Connection
- Select "Google Cloud Storage" from the Add Connection panel
-
Enter the necessary authentication properties to connect to Google Cloud Storage.
Authenticate with a User Account
You can connect without setting any connection properties for your user credentials. After setting InitiateOAuth to GETANDREFRESH, you are ready to connect.
When you connect, the Google Cloud Storage OAuth endpoint opens in your default browser. Log in and grant permissions, then the OAuth process completes
Authenticate with a Service Account
Service accounts have silent authentication, without user authentication in the browser. You can also use a service account to delegate enterprise-wide access scopes.
You need to create an OAuth application in this flow. See the Help documentation for more information. After setting the following connection properties, you are ready to connect:
- InitiateOAuth: Set this to GETANDREFRESH.
- OAuthJWTCertType: Set this to "PFXFILE".
- OAuthJWTCert: Set this to the path to the .p12 file you generated.
- OAuthJWTCertPassword: Set this to the password of the .p12 file.
- OAuthJWTCertSubject: Set this to "*" to pick the first certificate in the certificate store.
- OAuthJWTIssuer: In the service accounts section, click Manage Service Accounts and set this field to the email address displayed in the service account Id field.
- OAuthJWTSubject: Set this to your enterprise Id if your subject type is set to "enterprise" or your app user Id if your subject type is set to "user".
- ProjectId: Set this to the Id of the project you want to connect to.
The OAuth flow for a service account then completes.
- Click Save & Test
- Navigate to the Permissions tab and update the user-based permissions.

Generate a Personal Access Token (PAT)
The Python SDK authenticates to Connect AI with your account email and a Personal Access Token (PAT). It is best practice to create a separate PAT for each application to maintain granularity of access.
- Click the Gear icon () at the top right of the Connect AI app to open the Settings page.
- On the Settings page, go to the Access Tokens section and click Create PAT.
- Give the PAT a name and click Create.

- The PAT is only visible at creation, so copy it and store it securely.
Install the SDK
Install the SDK from PyPI with pip:
pip install cdata-connect-ai
Connect and Run Your First Query
Connect with your account email and PAT, then query sys_tables to discover every table available across your connected sources. Identifiers in Connect AI are three-part: <Connection>.<Schema>.<Table>, where the connection name defaults to the source name (for example, GoogleCloudStorage1).
import cdata_connect_ai
conn = cdata_connect_ai.connect(
username="[email protected]",
password="<your_pat>",
)
cur = conn.cursor()
# Discover what's available across your connected sources
cur.execute("SELECT CatalogName, SchemaName, TableName FROM sys_tables LIMIT 25")
for row in cur.fetchall():
print(row)
Pick any table from the results and query it directly:
cur.execute(
"SELECT Name, OwnerId "
"FROM [GoogleCloudStorage1].[GoogleCloudStorage].[Buckets] "
"LIMIT 10"
)
for row in cur.fetchall():
print(row)
Google Cloud Storage is a read-only source in Connect AI, so the SDK supports queries but not INSERT, UPDATE, or DELETE. Close the connection when you are finished:
conn.close()
That is the entire workflow: one package, a PAT, and standard DB-API calls. Because the SDK returns a normal DB-API connection, it drops straight into the rest of the Python data ecosystem. From here you can load Google Cloud Storage data into pandas, build ETL pipelines with petl, or power a Dash web app, all using this same connection.
More Information and Free Trial
Now you can query live Google Cloud Storage data from Python through the CData Connect AI Python SDK. For more information on connecting to Google Cloud Storage (and hundreds of other data sources), visit the Connect AI page. Sign up for a free trial and start working with live Google Cloud Storage data in Python.