How to work with Linear Data in Apache Spark using SQL
Apache Spark is a fast and general engine for large-scale data processing. When paired with the CData JDBC Driver for Linear, Spark can work with live Linear data. This article describes how to connect to and query Linear data from a Spark shell.
The CData JDBC Driver offers unmatched performance for interacting with live Linear data due to optimized data processing built into the driver. When you issue complex SQL queries to Linear, the driver pushes supported SQL operations, like filters and aggregations, directly to Linear and utilizes the embedded SQL engine to process unsupported operations (often SQL functions and JOIN operations) client-side. With built-in dynamic metadata querying, you can work with and analyze Linear data using native data types.
Install the CData JDBC Driver for Linear
Download the CData JDBC Driver for Linear installer, unzip the package, and run the JAR file to install the driver.
Start a Spark Shell and Connect to Linear Data
- Open a terminal and start the Spark shell with the CData JDBC Driver for Linear JAR file as the jars parameter:
$ spark-shell --jars /CData/CData JDBC Driver for Linear/lib/cdata.jdbc.linear.jar - With the shell running, you can connect to Linear with a JDBC URL and use the SQL Context load() function to read a table.
You can authenticate to Linear with a personal API key or with OAuth 2.0. The API key is the simplest option for connecting with your own Linear account.
Authenticating with an API Key
Set the following connection properties:
- AuthScheme: Set this to APIKey.
- APIKey: A Linear personal API key.
To create a personal API key, log in to Linear, open Settings > Security & access > Personal API keys, select New API key, and create it. Copy the key immediately, because Linear shows it only once.
Authenticating with OAuth
OAuth requires a custom OAuth application registered in Linear (Settings > API > OAuth applications), which provides the OAuthClientId and OAuthClientSecret. Two flows are supported:
- Authorization code: Set AuthScheme to OAuth, InitiateOAuth to GETANDREFRESH, and provide OAuthClientId, OAuthClientSecret, and the CallbackURL defined in your application (e.g., http://localhost:33333). The driver opens Linear in your browser so you can grant access.
- Client credentials: Set AuthScheme to OAuthClient and provide OAuthClientId and OAuthClientSecret. This authenticates the application itself, with no browser interaction, and suits machine-to-machine integrations.
By default, the driver requests the read,write scopes. The driver refreshes the access token automatically when it expires.
Built-in Connection String Designer
For assistance in constructing the JDBC URL, use the connection string designer built into the Linear JDBC Driver. Either double-click the JAR file or execute the jar file from the command-line.
java -jar cdata.jdbc.linear.jarFill in the connection properties and copy the connection string to the clipboard.
Configure the connection to Linear, using the connection string generated above.
scala> val linear_df = spark.sqlContext.read.format("jdbc").option("url", "jdbc:linear:AuthScheme=APIKey;APIKey=myAPIKey;").option("dbtable","Team").option("driver","cdata.jdbc.linear.LinearDriver").load() - Once you connect and the data is loaded you will see the table schema displayed.
Register the Linear data as a temporary table:
scala> linear_df.registerTable("team")-
Perform custom SQL queries against the Data using commands like the one below:
scala> linear_df.sqlContext.sql("SELECT id, name FROM Team WHERE key = ENG").collect.foreach(println)You will see the results displayed in the console, similar to the following:
Using the CData JDBC Driver for Linear in Apache Spark, you are able to perform fast and complex analytics on Linear data, combining the power and utility of Spark with your data. Download a free, 30 day trial of any of the hundreds of CData JDBC Drivers and get started today.