Connect and Query Live PingOne Data in Databricks with CData Connect AI
Databricks is a leading AI cloud-native platform that unifies data engineering, machine learning, and analytics at scale. Its powerful data lakehouse architecture combines the performance of data warehouses with the flexibility of data lakes. Integrating Databricks with CData Connect AI gives organizations live, real-time access to PingOne data without the need for complex ETL pipelines or data duplication—streamlining operations and reducing time-to-insights.
In this article, we'll walk through how to configure a secure, live connection from Databricks to PingOne using CData Connect AI. Once configured, you'll be able to access PingOne data directly from Databricks notebooks using standard SQL—enabling unified, real-time analytics across your data ecosystem.
Overview
Here is an overview of the simple steps:
- Step 1 — Connect and Configure: In CData Connect AI, create a connection to your PingOne source, configure user permissions, and generate a Personal Access Token (PAT).
- Step 2 — Query from Databricks: Install the CData JDBC driver in Databricks, configure your notebook with the connection details, and run SQL queries to access live PingOne data.
Prerequisites
Before you begin, make sure you have the following:
- An active PingOne account.
- A CData Connect AI account. You can log in or sign up for a free trial here.
- A Databricks account. Sign up or log in here.
Step 1: Connect and Configure a PingOne Connection in CData Connect AI
1.1 Add a Connection to PingOne
CData Connect AI uses a straightforward, point-and-click interface to connect to available data sources.
- Log into Connect AI, click Sources on the left, and then click Add Connection in the top-right.
- Select "PingOne" from the Add Connection panel.
-
Enter the necessary authentication properties to connect to PingOne.
To connect to PingOne, configure these properties:
- Region: The region where the data for your PingOne organization is being hosted.
- AuthScheme: The type of authentication to use when connecting to PingOne.
- Either WorkerAppEnvironmentId (required when using the default PingOne domain) or AuthorizationServerURL, configured as described below.
Configuring WorkerAppEnvironmentId
WorkerAppEnvironmentId is the ID of the PingOne environment in which your Worker application resides. This parameter is used only when the environment is using the default PingOne domain (auth.pingone). It is configured after you have created the custom OAuth application you will use to authenticate to PingOne, as described in Creating a Custom OAuth Application in the Help documentation.
First, find the value for this property:
- From the home page of your PingOne organization, move to the navigation sidebar and click Environments.
- Find the environment in which you have created your custom OAuth/Worker application (usually Administrators), and click Manage Environment. The environment's home page displays.
- In the environment's home page navigation sidebar, click Applications.
- Find your OAuth or Worker application details in the list.
-
Copy the value in the Environment ID field.
It should look similar to:
WorkerAppEnvironmentId='11e96fc7-aa4d-4a60-8196-9acf91424eca'
Now set WorkerAppEnvironmentId to the value of the Environment ID field.
Configuring AuthorizationServerURL
AuthorizationServerURL is the base URL of the PingOne authorization server for the environment where your application is located. This property is only used when you have set up a custom domain for the environment, as described in the PingOne platform API documentation. See Custom Domains.
Authenticating to PingOne with OAuth
PingOne supports both OAuth and OAuthClient authentication. In addition to performing the configuration steps described above, there are two more steps to complete to support OAuth or OAuthCliet authentication:
- Create and configure a custom OAuth application, as described in Creating a Custom OAuth Application in the Help documentation.
- To ensure that the driver can access the entities in Data Model, confirm that you have configured the correct roles for the admin user/worker application you will be using, as described in Administrator Roles in the Help documentation.
- Set the appropriate properties for the authscheme and authflow of your choice, as described in the following subsections.
OAuth (Authorization Code grant)
Set AuthScheme to OAuth.
Desktop Applications
Get and Refresh the OAuth Access Token
After setting the following, you are ready to connect:
- InitiateOAuth: GETANDREFRESH. To avoid the need to repeat the OAuth exchange and manually setting the OAuthAccessToken each time you connect, use InitiateOAuth.
- OAuthClientId: The Client ID you obtained when you created your custom OAuth application.
- OAuthClientSecret: The Client Secret you obtained when you created your custom OAuth application.
- CallbackURL: The redirect URI you defined when you registered your custom OAuth application. For example: https://localhost:3333
When you connect, the driver opens PingOne's OAuth endpoint in your default browser. Log in and grant permissions to the application. The driver then completes the OAuth process:
- The driver obtains an access token from PingOne and uses it to request data.
- The OAuth values are saved in the location specified in OAuthSettingsLocation, to be persisted across connections.
The driver refreshes the access token automatically when it expires.
For other OAuth methods, including Web Applications, Headless Machines, or Client Credentials Grant, refer to the Help documentation.
- Click Save & Test in the top-right.
-
Navigate to the Permissions tab on the PingOne Connection page
and update the user-based permissions based on your preferences.
1.2 Generate a Personal Access Token (PAT)
When connecting to Connect AI through the REST API, the OData API, or the Virtual SQL Server, a Personal Access Token (PAT) is used to authenticate the connection to Connect AI. PAT functions as an alternative to your login credentials for secure, token-based authentication. It is a best practice to create a separate PAT for each service to maintain granularity of access.
- Click on the Gear icon () at the top right of the Connect AI app to open the settings page.
- On the Settings page, go to the Access Tokens section and click Create PAT.
-
Give the PAT a name and click Create.
- Note: The personal access token is only visible at creation, so be sure to copy it and store it securely for future use.
Step 2: Connect and Query PingOne Data in Databricks
Follow these steps to establish a connection from Databricks to PingOne. You'll install the CData JDBC Driver for Connect AI, add the JAR file to your cluster, configure your notebooks, and run SQL queries to access live PingOne data data.
2.1 Install the CData JDBC Driver for Connect AI
- In CData Connect AI, click the Integrations page on the left. Search for JDBC or Databricks, click Download, and select the installer for your operating system.
-
Once downloaded, run the installer and follow the instructions:
- For Windows: Run the setup file and follow the installation wizard.
- For Mac/Linux: Unpack the archive and move the folder to /opt or /Applications. Make sure you have execute permissions.
-
After installation, locate the JAR file in the installation directory:
- Windows:
C:\Program Files\CData\CData JDBC Driver for Connect AI\lib\cdata.jdbc.connect.jar - Mac/Linux:
/Applications/CData/CData JDBC Driver for Connect AI/lib/cdata.jdbc.connect.jar
- Windows:
2.2 Install the JAR File on Databricks
-
Log in to Databricks. In the navigation pane, click Compute on the left. Start or create a compute cluster.
-
Click on the running cluster, go to the Libraries tab, and click Install New at the top right.
-
In the Install Library dialog, select DBFS, and drag and drop the
cdata.jdbc.connect.jar file. Click Install.
2.3 Query PingOne Data in a Databricks Notebook
Notebook Script 1 — Define JDBC Connection:
- Paste the following script into the notebook cell:
driver = "cdata.jdbc.connect.ConnectDriver"
url = "jdbc:connect:AuthScheme=Basic;User=your_username;Password=your_pat;URL=https://cloud.cdata.com/api/;DefaultCatalog=Your_Connection_Name;"
- Replace:
- your_username - With your CData Connect AI username
- your_pat - With your CData Connect AI Personal Access Token (PAT)
- Your_Connection_Name - With the name of your Connect AI data source, from the Sources page
- Run the script.
Notebook Script 2 — Load DataFrame from PingOne data:
- Add a new cell for this second script. From the menu on the right side of your notebook, click Add cell below.
- Paste the following script into the new cell:
remote_table = spark.read.format("jdbc") \
.option("driver", "cdata.jdbc.connect.ConnectDriver") \
.option("url", "jdbc:connect:AuthScheme=Basic;User=your_username;Password=your_pat;URL=https://cloud.cdata.com/api/;DefaultCatalog=Your_Connection_Name;") \
.option("dbtable", "YOUR_SCHEMA.YOUR_TABLE") \
.load()
- Replace:
- your_username - With your CData Connect AI username
- your_pat - With your CData Connect AI Personal Access Token (PAT)
- Your_Connection_Name - With the name of your Connect AI data source, from the Sources page
- YOUR_SCHEMA.YOUR_TABLE - With your schema and table, for example, PingOne.[CData].[Administrators].Users
- Run the script.
Notebook Script 3 — Preview Columns:
- Similarly, add a new cell for this third script.
- Paste the following script into the new cell:
display(remote_table.select("ColumnName1", "ColumnName2"))
- Replace ColumnName1 and ColumnName2 with the actual columns from your PingOne structure (e.g. Id, Username, etc.).
- Run the script.
You can now explore, join, and analyze live PingOne data directly within Databricks notebooks—without needing to know the complexities of the back-end API and without replicating PingOne data.
Try CData Connect AI Free for 14 Days
Ready to simplify real-time access to PingOne data? Start your free 14-day trial of CData Connect AI today and experience seamless, live connectivity from Databricks to PingOne.
Low code, zero infrastructure, zero replication — just seamless, secure access to your most critical data and insights.