DLT driver GC pressure during a large initial hydration. Is a bigger driver the only lever?
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
07-25-2026 01:56 PM
Last week one of our DLT pipelines spent five hours in RUNNING with nothing landing, and I'd like to compare notes on how people handle the underlying problem, because we found one fix and several dead ends.
Setup: A continuous DLT pipeline on classic compute (our secret access goes through a UC service credential, so serverless is not an option for this one). The pipeline is config-driven: 100+ source tables, each generating a CDC merge flow plus a historic load flow, so one update carries 200+ flows. In the environment with real volume, the initial hydration includes a 800 million row table.
The update reported RUNNING for hours, every flow RUNNING, zero batches committed. The pipeline event log told the real story: `gc_pressure` warnings every few minutes, with the driver paused for 60 plus seconds out of every 120 because of garbage collection. A driver paused half its life cannot schedule flows or commit anything. The takeaway that cost us a morning: RUNNING is a position, not progress, and the event log is where this failure is actually visible.
We had never set a `clusters` block, so the pipeline got the default cluster, whose driver is compute optimized with 16GB. The driver holds state for every flow in the update, and 200+ flows hydrating production volume at once did not fit. Notably, dev runs the identical 200+ flow graph on the same default driver without trouble. Flow count alone was fine; flow count multiplied by data scale was not.
Current fix: An explicit driver with double the memory and half the cores (which came out cheaper than the default, since the driver schedules work rather than crunching rows), set per environment, workers unchanged. GC warnings went to zero and the hydration completed. Checkpoints survived the stop and restart, so nothing was reprocessed.
What I'd like to learn from others,
1. Is there a supported way to cap how many flows hydrate concurrently within one update? We found `pipelines.maxConcurrentFlows` mentioned in a few threads, but it is not in the pipeline properties reference, so we dropped it rather than ship a key we couldn't substantiate. Does an official knob exist?
2. Are there rules of thumb for classic compute DLT driver sizing as flow count grows? The docs cover serverless sizing well; classic driver guidance is thin.
3. Would you have staged the backfill instead, hydrating subsets of tables across updates? How do you weigh that against one big update with a bigger driver?
4. Does anyone alert on `gc_pressure` events from the event log so this gets caught in minutes instead of hours?
- Labels:
-
Workflows