transformWithStateInPandas throws "Spark connect directory is not ready" error

felix4572
New Contributor III

Hello,

we employ arbitrary stateful aggregations in our data processing streams on Azure Databricks, and would like to migrate from applyInPandasWithState to transformWithStateInPandas. We employ the Python API throughout our solution, and some of our workspaces have NOT yet Unity Catalog enabled.

Trying to run the examples provided in the Azure Databricks documentation, e.g., the SCD Type 2 Example, on the workspaces without Unity Catalog enabled, I get the following error: 

felix4572_0-1756710186921.png

The cluster configuration is as follows:

  • DBR 17.1
  • Single node
  • Access mode "No isolation shared"
  • node type ID "Standard_D4ds_v5"
  • Photon not activated

To my understanding, this setup fullfils the requirements for using transformWithStateInPandas (DBR > 16.2, compute using "single user"/"dedicated" or "no isolation shared" access mode, using RocksDB as state store provider).

I also tested other examples, they all result in the same error when trying to start the stream. 

The exact same example with identical cluster configuration works in our Unity-enabled workspaces. 

What did I miss? Why is the spark connect directory not ready on the workspace that has Unity Catalog not enabled?

Best and thanks!

Felix