- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
04-02-2026 11:41 AM
Great observation. I can see why this feels like it contradicts the docs. It doesn’t, but I agree the documentation could be clearer about what’s happening under the hood. Let me walk through it.
Auto Loader’s schema inference actually operates in two layers, and that’s the key to what you’re seeing.
First, the schema file (stored in _schemas at your schemaLocation) always captures the actual detected types from sampling your data. That’s why Age shows up as a long there. Auto Loader needs those real types to track schema evolution over time, detect type changes, and decide what should land in _rescued_data when something doesn’t match.
Second, the DataFrame schema is where cloudFiles.inferColumnTypes comes into play. When that option is false, which is the default for JSON, CSV, and XML, Auto Loader takes the inferred schema and casts everything to strings before exposing it in the DataFrame. That’s the “safe default” the docs are referring to. When you flip it to true, the DataFrame reflects the actual detected types from the schema file instead of flattening everything to strings.
So when the docs say “all columns are inferred as strings,” they’re really talking about the DataFrame output, not the schema file itself.
You can see this clearly in your scenarios. In Scenario 1, the schema file correctly records Age as a long, but the DataFrame shows it as a string. In Scenario 2, once you enable inferColumnTypes, the DataFrame starts reflecting the real types. The schema file doesn’t change because it already had the correct types from the start.
Here’s the clean way to think about it:
Schema file
Always stores the true detected types, regardless of inferColumnTypes. This is by design for schema evolution.
DataFrame
Controlled by inferColumnTypes
False means everything is presented as strings
True means you get the actual detected types
Your instinct was spot on. Auto Loader is still inferring the data types, it just doesn’t always surface them in the DataFrame unless you tell it to.
Hope this helps, Lou.