Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
09-04-2025 11:25 PM
Hi @saicharandeepb ,
The behaviour you're experiencing can happen with coalesce. The thing is, when you use coalesce(1), you're sacrificing parallelism and everything is performed on a single executor.
There's even a warning in Apache Spark OSS regarding this:
You can also check following posts/blogs:
apache spark - does coalesce(1) the dataframe before write have any impact on performance? - Stack O...
(22) Analyzing a 30x Slowdown in My Spark Program Due to Coalesce | LinkedIn