โ03-28-2026 09:41 AM
Hi everyone, Iโm facing an issue when running a notebook on a Databricks All-purpose cluster. Some of my cells/pipelines run for a very long time, and I want to automatically cancel/kill them when they exceed a certain time limit.
I tried setting spark.databricks.execution.timeout, but it doesnโt seem to have any effect in my case.
What I need is a timeout mechanism that can cancel the currently running notebook cell, not just a Spark job timeout.
If anyone can share guidance or official documentation references, Iโd really appreciate it. Thanks in advance!
โ03-28-2026 11:25 AM
#Add in notebook
import signal
class TimeoutException(Exception):
"""Raised when a cell is run for very long time"""
def timeout_handler(signum, frame):
raise TimeoutException("Timed out!")
def set_cell_timeout(seconds):
signal.signal(signal.SIGALRM, timeout_handler)
signal.alarm(seconds)
#Add in a notebook cell running notebook function
try:
set_cell_timeout(30) # Set for 30 seconds
#notebook function
finally:
signal.alarm(0)โ03-28-2026 07:20 PM
What Iโm looking for is a workspace-level monitoring approach: detect any notebook execution where a cell (or the run) has been running longer than a threshold, and then cancel/terminate it automatically.
Iโve tried looking into audit tables, REST APIs, but it seems they donโt provide enough visibility at cell-level
โ03-28-2026 09:58 PM
For the issue - Some of my cells/pipelines run for a very long time, and I want to automatically cancel/kill them when they exceed a certain time limit.
โ03-29-2026 07:09 PM
@zenwanderer Have you looked into Query Watchdog?
For Classic All-Purpose clusters this might be your best bet.
https://docs.databricks.com/aws/en/compute/troubleshooting/query-watchdog
Monday
Greetings @zenwanderer, I did some digging and here is what I found.
Your follow-up narrows the ask: you want the workspace, not each developer, to catch and kill any interactive cell that runs past a threshold. Short answer: there's no supported, built-in mechanism that does that on all-purpose compute today. I also couldn't find spark.databricks.execution.timeout in the docs, which is likely why it had no effect.
Here's how the existing controls line up, since they each work at a different scope:
signal.alarm snippet is a fine best-effort guard around Python code in an interactive notebook. It's opt-in per cell and won't reliably stop Spark or JVM work, so treat it as a convenience, not a control.On the monitoring side you already explored: verbose audit logs do carry cell-level data (notebook / runCommand with commandId, notebookId, and executionTime), but the event fires after the command finishes, so it's good for reporting on long-running cells, not for killing them in flight. The Command Execution API can cancel commands, but only in contexts your own code created through the API, not cells launched from the UI.
If you still want in-flight cancellation on a classic all-purpose cluster, the workaround I'd test is a driver-side watchdog: a background thread started at cluster launch that polls the driver's Spark REST endpoint (via spark.sparkContext.uiWebUrl) for active jobs and calls sc.cancelJobGroup(...) on any job group older than your threshold. Each notebook command gets its own job group, so cancelling it fails the cell. It only stops Spark work, though, and it's a workaround rather than a supported feature, so try it on a dev cluster first.
At the end of the day, the pattern that holds up is interactive notebooks with Query Watchdog and auto-termination as guardrails, and anything that resembles a pipeline moved to a job with a timeout.
References:
Regards,
Louis