Repository navigation
Replies: 1 comment
|
Hey @Somesh, Good question — this is one of the biggest points of confusion with the Databricks Jobs API.
Here's the practical difference: If you use jobs.SubmitTask(
task_key='my_task',
existing_cluster_id='cluster-abc123',
notebook_task=jobs.NotebookTask(notebook_path='/my/notebook')
)Your job runs immediately (no wait for cluster startup), but you're paying for that cluster sitting idle when nothing's running. If you use job_clusters=[
JobCluster(
job_cluster_key="default",
new_cluster={
"spark_version": "14.0.x-scala2.12",
"node_type_id": "i3.xlarge",
"num_workers": 2
}
)
],
tasks=[
Task(
task_key="my_task",
job_cluster_key="default",
notebook_task=NotebookTask(notebook_path="/my/notebook")
)
]Your job takes 5-10 minutes to start (cluster spinup time), but the cluster gets deleted automatically when done, so you only pay for what you use. For your parallel notebook runner scenario in that thread you linked — you're doing the right thing using If you were doing one-off jobs, Does that help clarify? |
Uh oh!
There was an error while loading. Please reload this page.
We currently have a notebook that executes other notebooks in parallel (number of notebooks depends on runtime parameters).
The current solution is to execute
dbutils.notebooks.runwithinThreadPoolExecutor. I'm worried that this might not make optimal use of Databricks cluster resources and would like to use native scheduling functionality.In the example within this repo the JobsAPI is used to execute a dynamically generated job, which seems nice.
My question is the following: How can I
I essentially want to execute the tasks in parallel, wait until all of them are either finished, cancelled or failed and then gather exceptions if any were raised and further deal with them in the caller notebook.
I am now using this bit of code in my caller notebook to run some tests:
The notebook that is called contains
All reactions