- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
04-01-2026 10:58 AM - edited 04-01-2026 11:21 AM
Hi @Louis_Frolio ,
I explored something related and intresteing (or confusing).
This conflicts with Databrick's documentation statement as follows
"By default, Auto Loader schema inference seeks to avoid schema evolution issues due to type mismatches. For formats that don't encode data types (JSON, CSV, and XML), Auto Loader infers all columns as strings (including nested fields in JSON files)."
What I experinced recenlty is that Autoloader DOES inefer schema for json file IN the SCHEMA FILE it creates at the schema location.Also,it infers all columns as string IN THE DATAFRAME ONLY.
Below is my observation.
Input Data :
{"Name":"Alfred","geneder":"M","Age":14}
{"Name":"John","geneder":"M","Age":12}
Scenario 1 : Without cloudFiles.inferColumnTypes
{"dataSchemaJson":"{\"type\":\"struct\",\"fields\":[{\"name\":\"Age\",\"type\":\"long\",\"nullable\":true,\"metadata\":{}},{\"name\":\"Name\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}},{\"name\":\"geneder\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}}]}","partitionSchemaJson":"{\"type\":\"struct\",\"fields\":[]}"}
{"dataSchemaJson":"{\"type\":\"struct\",\"fields\":[{\"name\":\"Age\",\"type\":\"long\",\"nullable\":true,\"metadata\":{}},{\"name\":\"Name\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}},{\"name\":\"geneder\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}}]}","partitionSchemaJson":"{\"type\":\"struct\",\"fields\":[]}"}