Databricks Community

Jasam · ‎07-19-2016

I am using spark- csv utility, but I need when it infer schema all columns be transform in string columns by default.

Thanks in advance.

User16789201666 · ‎07-22-2016

You can manually specify schema, e.g. from (https://github.com/databricks/spark-csv):

import org.apache.spark.sql.SQLContext import org.apache.spark.sql.types.{StructType, StructField, StringType, IntegerType};

val sqlContext = new SQLContext(sc) val customSchema = StructType(Array( StructField("year", IntegerType, true), StructField("make", StringType, true), StructField("model", StringType, true), StructField("comment", StringType, true), StructField("blank", StringType, true)))

val df = sqlContext.read .format("com.databricks.spark.csv") .option("header", "true") // Use first line of all files as header .schema(customSchema) .load("cars.csv")

val selectedData = df.select("year", "model") selectedData.write .format("com.databricks.spark.csv") .option("header", "true") .save("newcars.csv")

vadeka · ‎11-15-2018

I was solving the same issue, that I wanted all the columns as text and deal with correct cast later which I have solved by recasting all the column to string after I've inferred the schema. I'm not sure if it's efficient, but it works.

#pyspark path = '...' df = spark.read \ .option("inferschema", "true") \ .csv(df)

for column in df.columns: df= df.withColumn(column,df[column].cast('string'))

then you have to read again with changed schema

f = spark.read.option("schema", df.schema).csv(df)

This however doesn't deal with nested columns, though csv doesn't create any nested structs, I hope.