Kill/Cancel a Notebook Cell Running Too Long on an All-purpose Cluster
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-28-2026 09:41 AM
Hi everyone, I’m facing an issue when running a notebook on a Databricks All-purpose cluster. Some of my cells/pipelines run for a very long time, and I want to automatically cancel/kill them when they exceed a certain time limit.
I tried setting spark.databricks.execution.timeout, but it doesn’t seem to have any effect in my case.
What I need is a timeout mechanism that can cancel the currently running notebook cell, not just a Spark job timeout.
If anyone can share guidance or official documentation references, I’d really appreciate it. Thanks in advance!
- Labels:
-
Spark
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-28-2026 11:25 AM
- You can use signal to do this if running in a notebook for code validation
#Add in notebook
import signal
class TimeoutException(Exception):
"""Raised when a cell is run for very long time"""
def timeout_handler(signum, frame):
raise TimeoutException("Timed out!")
def set_cell_timeout(seconds):
signal.signal(signal.SIGALRM, timeout_handler)
signal.alarm(seconds)
#Add in a notebook cell running notebook function
try:
set_cell_timeout(30) # Set for 30 seconds
#notebook function
finally:
signal.alarm(0)- You can use lakeflow job notifications with threshold to cancel jobs running too long. Avoid using signals in notebooks running via jobs
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-28-2026 07:20 PM
What I’m looking for is a workspace-level monitoring approach: detect any notebook execution where a cell (or the run) has been running longer than a threshold, and then cancel/terminate it automatically.
I’ve tried looking into audit tables, REST APIs, but it seems they don’t provide enough visibility at cell-level
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-28-2026 09:58 PM
For the issue - Some of my cells/pipelines run for a very long time, and I want to automatically cancel/kill them when they exceed a certain time limit.
- You can use job notifications with Metric threshold (Duration Warning for notifications & Duration Timeout for kill) to cancel jobs running too long (completion time more than Duration Timeout). More details here
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
03-29-2026 07:09 PM
@zenwanderer Have you looked into Query Watchdog?
For Classic All-Purpose clusters this might be your best bet.
https://docs.databricks.com/aws/en/compute/troubleshooting/query-watchdog
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
Monday
Greetings @zenwanderer, I did some digging and here is what I found.
Your follow-up narrows the ask: you want the workspace, not each developer, to catch and kill any interactive cell that runs past a threshold. Short answer: there's no supported, built-in mechanism that does that on all-purpose compute today. I also couldn't find spark.databricks.execution.timeout in the docs, which is likely why it had no effect.
Here's how the existing controls line up, since they each work at a different scope:
- @balajij8's
signal.alarmsnippet is a fine best-effort guard around Python code in an interactive notebook. It's opt-in per cell and won't reliably stop Spark or JVM work, so treat it as a convenience, not a control. - His second point is the one that matters most: if these are pipelines, they belong in Lakeflow Jobs. Set a timeout under the task's Metric thresholds (Run duration, then Timeout) or at the job level, and Databricks marks the run Timed Out when it overruns. That's the only real "kill it after N minutes" switch the platform gives you, and it applies to job runs, not to someone manually running a cell.
- @MoJaMa's Query Watchdog is the right guardrail for shared all-purpose clusters, with one nuance: it isn't a time limit. It cancels tasks whose output rows blow past a ratio of input rows (default 1000x) or that fan out into too many tasks or partitions. It'll catch a skewed join that would never finish, but not a query that's slow for legitimate reasons. Set it in the cluster Spark config so it covers everyone.
On the monitoring side you already explored: verbose audit logs do carry cell-level data (notebook / runCommand with commandId, notebookId, and executionTime), but the event fires after the command finishes, so it's good for reporting on long-running cells, not for killing them in flight. The Command Execution API can cancel commands, but only in contexts your own code created through the API, not cells launched from the UI.
If you still want in-flight cancellation on a classic all-purpose cluster, the workaround I'd test is a driver-side watchdog: a background thread started at cluster launch that polls the driver's Spark REST endpoint (via spark.sparkContext.uiWebUrl) for active jobs and calls sc.cancelJobGroup(...) on any job group older than your threshold. Each notebook command gets its own job group, so cancelling it fails the cell. It only stops Spark work, though, and it's a workaround rather than a supported feature, so try it on a dev cluster first.
At the end of the day, the pattern that holds up is interactive notebooks with Query Watchdog and auto-termination as guardrails, and anything that resembles a pipeline moved to a job with a timeout.
References:
- Query Watchdog: https://docs.databricks.com/aws/en/compute/troubleshooting/query-watchdog
- Configure jobs: https://docs.databricks.com/aws/en/jobs/configure-job
- Task duration thresholds: https://learn.microsoft.com/en-us/azure/databricks/jobs/configure-task#configure-thresholds-for-task...
- Verbose audit logs: https://docs.databricks.com/aws/en/admin/account-settings/verbose-logs
Regards,
Louis