Autoloader issue
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
11-21-2023 05:57 AM
I'm trying to ingest data from Parquet files using Autoloader. Now, I have my custom schema, I don't want to infer the schema from the parquet files.
During readstream everything is fine. But during writestream, it is somehow inferring the schema from the files and I'm getting a schema mismatch error.
Any idea why it is happening? Help will be appreciated.
#
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
12-05-2024 09:22 AM
In this case, please make sure you specify the schema explicitly when reading the Parquet files and do not specify any inference options.
Something like
spark.readStream.format("cloudFiles").schema(schema)...
If you want to more easily grab the schema, you can read with the batch reader and capture the schema:
schema = spark.read.parquet("/your/path/here").schema