CREATE JOB
Description
Create a new Job.
Note that the SQL statement does not end with ;
Syntax
CREATE JOB `<job_name>`
TYPE = { 'JAR' | 'PYTHON' }
PARAMETERS = <array>
CLUSTER = <string>
[ MAX_CONCURRENT_RUNS = <integer> ]
Example
Example command for creating a Job.
CREATE JOB `count_transactions`
TYPE = 'JAR'
PARAMETERS = ( '--class', 'com.example.MySparkApp', '/path/to/my-spark-app.jar', 'arg1', 'arg2' )
CLUSTER = 'onehouse_cluster_spark'
Example command for creating a Job that allows up to 3 runs at the same time.
CREATE JOB `count_transactions`
TYPE = 'JAR'
PARAMETERS = ( '--class', 'com.example.MySparkApp', '/path/to/my-spark-app.jar', 'arg1', 'arg2' )
CLUSTER = 'onehouse_cluster_spark'
MAX_CONCURRENT_RUNS = 3
Required parameters
<job_name>: Unique name to identify the Job (max 100 characters).TYPE: Specify the type of Job - this can be a JAR (for Java or Scala code) or Python script.PARAMETERS: Specify an array of Strings to pass as parameters to the Job, which will be used in aspark-submit(see Apache Spark docs). This should include the following:- [Required] For JAR Jobs, you must include the
--classparameter. - [Required] Include the cloud storage bucket path containing the code for your Job. The Onehouse agent must have access to read this path.
- [Optional] Include any other Spark properties you'd like the Job to use.
- [Optional] Include Apache Hudi configurations as
--hudi-conf <key>=<value>pairs — see Set Apache Hudi configurations on a Job. - [Optional] Include any arguments you'd like to pass to the Job.
- [Required] For JAR Jobs, you must include the
CLUSTER: Specify the name of an existing Onehouse Cluster with typeSparkto run the Job.
Optional parameters
MAX_CONCURRENT_RUNS: Maximum number of runs of this Job that can be active (Queued or Running) at the same time. Defaults to1, which is the behavior when the clause is omitted. See Concurrent Job runs for the requirements and limits.- Must be a whole number
>= 1. - Any value greater than
1requires concurrent runs to be enabled for your project — contact Onehouse support. Without it, the command fails withConcurrent runs are not enabled for this organization; max_concurrent_runs must be 1. - Each concurrent run is a separate Spark driver on the Cluster, so size the Cluster for the number of runs you allow.
- Must be a whole number
Status API
Status API response
API_OPERATION_STATUS_SUCCESSfrom Status API indicates that the Job has been created.API_OPERATION_STATUS_FAILEDfrom Status API does not necessarily mean the Job was not created. To confirm whether the Job was created, send a request to DESCRIBE JOB API.
Example Status API response
The Status API response of a successful Job creation.
{
"apiStatus": "API_OPERATION_STATUS_SUCCESS",
"apiResponse": {}
}