Spark issue handling data from json when the schema DataType mismatch occurs
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
01-05-2023 06:57 AM
Hi,
I have encountered a problem using spark, when creating a dataframe from a raw json source.
I have defined an schema for my data and the problem is that when there is a mismatch between one of the column values and its defined schema, spark not only sets that column to null but also sets all the columns after that column to null. I made a minimal example to show the problem as attached.
Basically, the values of "foo71" and "foo72" for the row_2 have issue with the defined schema "DecimalType(10,7)". Then, in the created dataframe (df_1), all the column after the column "foo7" return null, despite the schema is correct for those. What actually I expect is that only the "foo72" and "foo71" to be nulls, and the rest of columns after that should have the correct value in the columns.
I have used the following for this notebook:
1- Databricks Notebook in the databricks data science and engineering workspace
2- My cluster uses Databricks Runtime Version 12.0 (includes Apache Spark 3.3.1, Scala 2.12)
I have a workaround this problem, but I thought that it might worth reporting it to have a fix, because this can potentially risks losing data in similar situations.
Thanks you
Farzad
- Labels:
-
Apache spark
-
Column Values