juanlozadab
New Contributor III

I'll answer the size question directly first: no, we don't have a GB threshold, and I've come to think a number there doesn't survive contact with reality.

The reason is that the same 5 GB behaves completely differently depending on shape. Five gigabytes filtered and written out is trivially a single-machine job. Five gigabytes going through a wide join against a dimension table, or a window function over a skewed key, is not. And no threshold expressed in gigabytes tells you which one you're holding. So a size rule ends up either so conservative it never fires, or it fires on the job that turns out to need the shuffle.

What I'd threshold on instead is where the wall clock actually goes. Most "Spark is overkill" complaints I read are startup complaints, and those have a config fix rather than an architecture fix. Serverless jobs in standard mode carry a documented 4-6 minute startup latency by design. Performance optimized mode exists for exactly this case, and instance pools do the same on classic compute. If a 3 GB pipeline takes 7 minutes and 5 are provisioning, a second engine fixes the 2 minutes nobody was complaining about.

Which also answers your single node question, I think: single node removes the shuffle, not the cold start. So it helps if your bottleneck was coordination overhead, and does nothing if your bottleneck was waiting for compute. Worth knowing which one you have before picking.

On the second engine, what decides it for me is the asymmetry. A small job on Spark is a tax: bounded, predictable, shrinking every time start times improve. A job that outgrew the small engine is a rewrite: unbounded, and it lands at the worst possible moment. I'll take a known tax over an unknown rewrite. And the cost people underestimate isn't the tooling, it's that a second execution path doubles the failure modes. That's an on-call conversation, not a performance one.

So yes, a lot of this is the cost of standardisation. I don't think that's a cop-out, just something worth naming rather than dressing up as a performance decision.

 

jlb