Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
02-05-2026 01:58 PM
Does df.printSchema show region and days as partition columns at the end? If not, partition discovery isn’t working.
Can you remove mergeSchema or provide an explicit schema? - With mergeSchema, Spark must read the Parquet footers of all files under the base path to merge column definitions before planning the scan. This happens prior to partition pruning, so you’ll see a read/list of the full tree even if the final scan prunes most files.
Thank You
Pradeep Singh - https://www.linkedin.com/in/dbxdev
Pradeep Singh - https://www.linkedin.com/in/dbxdev