yesterday
Hi everyone,
A question I keep having with my team and I would like to hear how others think about it.
A lot of the jobs we run are not big. Plenty of our pipelines process a few gigabytes, some considerably less. We run them on Spark because that is what the platform is, but a single machine with a columnar engine would handle most of them comfortably, and often faster, because there is no shuffle and no cluster to start.
The counterargument is real too. One engine means one set of skills, one deployment pattern, one place to look when something breaks. Introducing a second execution path for small jobs means maintaining two of everything, and the small job that quietly grows into a large one becomes a migration problem.
So I am curious where people have landed.
Do you have a size threshold below which you deliberately avoid Spark, or do you run everything through it regardless?
For those running small jobs on Spark, is single node compute the answer, or do you find the overhead still dominates?
Has anyone introduced a second engine for small workloads and later regretted it? Or regretted not doing it?
And the honest version of the question: how much of this is a real engineering decision and how much is just the cost of standardisation being worth paying?
Interested in how teams reasoned about it, not just what they picked.
17 hours ago
Hi @Islam_hoti ,
This is the 'Silent Architect' dilemmaโwhen the standard tool is technically suboptimal but organizationally necessary. Weโve wrestled with this, and hereโs where we landed:
Our Verdict: We now run everything through Spark regardless of size. The 'Cost of Standardization' is an insurance policy against migration debt. If the job is truly tiny, we optimize the compute (Serverless/Single-node) rather than changing the engine.
Itโs an engineering decision disguised as a cost one: standardizing on one engine creates 'velocity' that often outweighs 'efficiency' in the long run.
yesterday
I'll answer the size question directly first: no, we don't have a GB threshold, and I've come to think a number there doesn't survive contact with reality.
The reason is that the same 5 GB behaves completely differently depending on shape. Five gigabytes filtered and written out is trivially a single-machine job. Five gigabytes going through a wide join against a dimension table, or a window function over a skewed key, is not. And no threshold expressed in gigabytes tells you which one you're holding. So a size rule ends up either so conservative it never fires, or it fires on the job that turns out to need the shuffle.
What I'd threshold on instead is where the wall clock actually goes. Most "Spark is overkill" complaints I read are startup complaints, and those have a config fix rather than an architecture fix. Serverless jobs in standard mode carry a documented 4-6 minute startup latency by design. Performance optimized mode exists for exactly this case, and instance pools do the same on classic compute. If a 3 GB pipeline takes 7 minutes and 5 are provisioning, a second engine fixes the 2 minutes nobody was complaining about.
Which also answers your single node question, I think: single node removes the shuffle, not the cold start. So it helps if your bottleneck was coordination overhead, and does nothing if your bottleneck was waiting for compute. Worth knowing which one you have before picking.
On the second engine, what decides it for me is the asymmetry. A small job on Spark is a tax: bounded, predictable, shrinking every time start times improve. A job that outgrew the small engine is a rewrite: unbounded, and it lands at the worst possible moment. I'll take a known tax over an unknown rewrite. And the cost people underestimate isn't the tooling, it's that a second execution path doubles the failure modes. That's an on-call conversation, not a performance one.
So yes, a lot of this is the cost of standardisation. I don't think that's a cop-out, just something worth naming rather than dressing up as a performance decision.
19 hours ago
Most teams ultimately stick with Spark for small workloads despite the overhead, because the hidden costs of maintaining a dual-engine stackโfragmented skill sets, separate deployment pipelines, and migration headaches when small jobs scaleโusually outweigh raw performance gains. However, if the platform overhead routinely stalls velocity or balloons cloud costs, introducing a lightweight single-node alternative like DuckDB or Polars is justified, provided your team has the operational capacity to manage two distinct execution paths.
17 hours ago
Hi @Islam_hoti ,
This is the 'Silent Architect' dilemmaโwhen the standard tool is technically suboptimal but organizationally necessary. Weโve wrestled with this, and hereโs where we landed:
Our Verdict: We now run everything through Spark regardless of size. The 'Cost of Standardization' is an insurance policy against migration debt. If the job is truly tiny, we optimize the compute (Serverless/Single-node) rather than changing the engine.
Itโs an engineering decision disguised as a cost one: standardizing on one engine creates 'velocity' that often outweighs 'efficiency' in the long run.