Databricks runs cell, but stops output and hangs afterwards.

ThomasKastl
Contributor

tl;dr: A cell that executes purely on the head node stops printed output during execution, but output still shows up in the cluster logs. After execution of the cell, Databricks does not notice the cell is finished and gets stuck. When trying to cancel, Databricks gets stuck as well, and we need to "Clear state".

Long version:

We use the tsfresh library (https://github.com/blue-yonder/tsfresh) in Databricks on a head node (no Spark - just Python). On most runs, the output of the notebook cell simply stops - while the cell is still being executed. This means that in the notebook itself, no new output is shown, even though the cell keeps running in the background. We know this because files generated by this cell are still written, and also, in Cluster -> Driver Logs, output keeps appearing.

This in itself wouldn't really be a problem, however, Databricks doesn't ever realize the cell is finished - meaning the next cell never gets executed. Also, the cell cannot be cancelled the regular way, we need to clear state, meaning losing all computation results that haven't been written out. Simply cancelling gets stuck.

This happened with Runtime 7.3 LTS, we switched to 10.4 LTS now and the problem is still persists. We tried different head node sizes and sometimes it gets stuck sooner, sometimes later, the behavior isn't consistent. We assume it has something to do with how tsfresh handles multitasking, but the problem seems to happen even if we turn off multitasking.

On local versions of Python notebooks, this never happens, leading us to assume it is a problem / bug with Databricks itself.

Any pointers what we can try / how we get in contact with someone from Databricks to check this?