mits1
New Contributor III

Hi @Louis_Frolio ,

I explored something related and intresteing (or confusing).

This conflicts with Databrick's documentation statement as follows

"By default, Auto Loader schema inference seeks to avoid schema evolution issues due to type mismatches. For formats that don't encode data types (JSON, CSV, and XML), Auto Loader infers all columns as strings (including nested fields in JSON files)."

What I experinced recenlty is that Autoloader DOES inefer schema for json file IN the SCHEMA FILE it creates at the schema location.Also,it infers all columns as string IN THE DATAFRAME ONLY.

Below is my observation.

Input Data :

{"Name":"Alfred","geneder":"M","Age":14}
{"Name":"John","geneder":"M","Age":12}

Scenario 1 : Without cloudFiles.inferColumnTypes

df = spark.readStream.\
    format("cloudFiles")\
    .option("cloudFiles.format", "json")\
    .option("cloudFiles.schemaLocation", "/Volumes/workspace/default/sys/schema4")\
    .load('/Volumes/workspace/dev/input/')
 
DF schema -> 
Age:string
Name:string
geneder:string
_rescued_data:string
 
Autoloader schema file contents ->
v1
{"dataSchemaJson":"{\"type\":\"struct\",\"fields\":[{\"name\":\"Age\",\"type\":\"long\",\"nullable\":true,\"metadata\":{}},{\"name\":\"Name\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}},{\"name\":\"geneder\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}}]}","partitionSchemaJson":"{\"type\":\"struct\",\"fields\":[]}"}
 
Note :- Looks like autoloader still infers the data but not the dataframe.
 ---------------------------------------------------------------------------------------------------------------------------------------------------------
Scenario 2 : With cloudFiles.inferColumnTypes
Input data : 
{"Name":"Mits","geneder":"F","Age":35}
 
df = spark.readStream.\
    format("cloudFiles")\
    .option("cloudFiles.format", "json")\
    .option("cloudFiles.inferColumnTypes", "true")\
    .option("cloudFiles.schemaLocation", "/Volumes/workspace/default/sys/schema4")\
    .load('/Volumes/workspace/dev/input/')
 
DF Schema ->
 
Age:long
Name:string
geneder:string
_rescued_data:string
 
Autoloader schema file contents -> No change!!
 
v1
{"dataSchemaJson":"{\"type\":\"struct\",\"fields\":[{\"name\":\"Age\",\"type\":\"long\",\"nullable\":true,\"metadata\":{}},{\"name\":\"Name\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}},{\"name\":\"geneder\",\"type\":\"string\",\"nullable\":true,\"metadata\":{}}]}","partitionSchemaJson":"{\"type\":\"struct\",\"fields\":[]}"}
 
Note : Dataframe's schema changes but not the Autoloader's (obviosuly,because there is no change in the source data).
 
It would be very helpful for me to understand behaviour of schema inference,if you could clarify this.
Thanks in advance!!