How to work with UKG Pro HCM Data in Apache Spark using SQL
Apache Spark is a fast and general engine for large-scale data processing. When paired with the CData JDBC Driver for UKG Pro HCM, Spark can work with live UKG Pro HCM data. This article describes how to connect to and query UKG Pro HCM data from a Spark shell.
The CData JDBC Driver offers unmatched performance for interacting with live UKG Pro HCM data due to optimized data processing built into the driver. When you issue complex SQL queries to UKG Pro HCM, the driver pushes supported SQL operations, like filters and aggregations, directly to UKG Pro HCM and utilizes the embedded SQL engine to process unsupported operations (often SQL functions and JOIN operations) client-side. With built-in dynamic metadata querying, you can work with and analyze UKG Pro HCM data using native data types.
Install the CData JDBC Driver for UKG Pro HCM
Download the CData JDBC Driver for UKG Pro HCM installer, unzip the package, and run the JAR file to install the driver.
Start a Spark Shell and Connect to UKG Pro HCM Data
- Open a terminal and start the Spark shell with the CData JDBC Driver for UKG Pro HCM JAR file as the jars parameter:
$ spark-shell --jars /CData/CData JDBC Driver for UKG Pro HCM/lib/cdata.jdbc.api.jar - With the shell running, you can connect to UKG Pro HCM with a JDBC URL and use the SQL Context load() function to read a table.
Start by setting the Profile connection property to the location of the UKGProHCM Profile on disk (e.g. C:\profiles\UKGProHCM.apip). Next, set the ProfileSettings connection property to the connection string for UKGProHCM (see below).
UKGProHCM API Profile Settings
UKG Pro HCM uses Basic authentication combined with a customer API key to authorize access to the API. To connect, you will need your UKG Pro HCM credentials and a customer API key issued by UKG.
Set the following connection properties to authenticate:
- AuthScheme: Set this to Basic.
- User: Set this to your UKG Pro HCM username.
- Password: Set this to your UKG Pro HCM password.
- ProfileSettings: Set this to a semicolon-separated list containing:
- Host: Your UKG Pro HCM tenant hostname (e.g., yourcompany.ultipro.com).
- CustomerAPIKey: Your UKG customer API key. This is passed as the US-CUSTOMER-API-KEY request header on each API call.
To obtain your customer API key, contact your UKG Pro HCM administrator or refer to the UKG Developer Portal at https://developer.ukg.com.
Built-in Connection String Designer
For assistance in constructing the JDBC URL, use the connection string designer built into the UKG Pro HCM JDBC Driver. Either double-click the JAR file or execute the jar file from the command-line.
java -jar cdata.jdbc.api.jarFill in the connection properties and copy the connection string to the clipboard.
Configure the connection to UKG Pro HCM, using the connection string generated above.
scala> val api_df = spark.sqlContext.read.format("jdbc").option("url", "jdbc:api:Profile=C:\profiles\UKGProHCM.apip;AuthScheme=Basic;User=your_username;Password=your_password;ProfileSettings='Host=yourcompany.ultipro.com;CustomerAPIKey=your_customer_api_key';").option("dbtable","UserDetails").option("driver","cdata.jdbc.api.APIDriver").load() - Once you connect and the data is loaded you will see the table schema displayed.
Register the UKG Pro HCM data as a temporary table:
scala> api_df.registerTable("userdetails")-
Perform custom SQL queries against the Data using commands like the one below:
scala> api_df.sqlContext.sql("SELECT UserId, UserName FROM UserDetails WHERE UserStatus = A").collect.foreach(println)You will see the results displayed in the console, similar to the following:
Using the CData JDBC Driver for UKG Pro HCM in Apache Spark, you are able to perform fast and complex analytics on UKG Pro HCM data, combining the power and utility of Spark with your data. Download a free, 30 day trial of any of the hundreds of CData JDBC Drivers and get started today.