- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
09-29-2025 10:05 AM
Hey @JaydeepKhatri here are some helpful points to consider:
Is this an officially supported, enhanced feature of the Databricks CSV reader?
-
Based on internal research, this appears to be an undocumented “feature” of Spark running on Databricks. Anecdotally, people are using it without major issues, but since it isn’t officially documented, proceed with caution.
Is this the recommended best practice on Databricks for ingesting CSVs with inconsistent schemas?
-
A better approach is to use Auto Loader, which reads files from cloud storage and provides schema evolution options. Specifically, look into the schemaEvolutionMode option.
-
Documentation: Schema evolution with Auto Loader
Example code:
df = (spark.readStream
.format("cloudFiles")
.option("cloudFiles.format", "csv")
.option("cloudFiles.schemaEvolutionMode", "addNewColumns")
.load("s3://path/to/files"))
Key distinctions between mergeSchema and Auto Loader schema evolution:
-
mergeSchema → A Delta Lake option used during reads/writes to Delta tables.
-
Auto Loader schema evolution → Configuration options (cloudFiles.schemaEvolutionMode) for handling new columns while streaming in CSV/JSON/Parquet.
👉 Direct answer: mergeSchema is not part of the Auto Loader API. Auto Loader has its own schema evolution features. However, if you’re ultimately writing the ingested data to Delta tables, you can use them together:
-
Auto Loader detects schema drift and evolves the DataFrame schema.
-
mergeSchema ensures the Delta table evolves when the data is written.
Hope this helps, Louis.