szymon_dybczak
Esteemed Contributor III

Hi @saicharandeepb ,

The behaviour you're experiencing can happen with coalesce. The thing is, when you use coalesce(1), you're sacrificing parallelism and everything is performed on a single executor.

There's even a warning in Apache Spark OSS regarding this:

szymon_dybczak_0-1757053069172.png

You can also check following posts/blogs:

apache spark - does coalesce(1) the dataframe before write have any impact on performance? - Stack O...

(22) Analyzing a 30x Slowdown in My Spark Program Due to Coalesce | LinkedIn