Witold
Databricks Partner

Why don't you simply use spark to process it? Like:

 

df = spark.read.option('sep','|').option('header', True).format('csv').load(file_path)

 

 Since it appears that you also have a schema you can avoid inferring it and pass it explicitly:

 

df = spark.read.option('sep','|').option('header', True).schema('PAT_KEY INT, CD_VERSION INT, CD_CODE STRING, CD_PRI_SEC STRING, CD_POA STRING').format('csv').load(file_path)

 

Then you process it and write it wherever and in which format you prefer.

If some of your data is corrupted, i.e. not according to the schema, you might want to look into Auto Loader and its rescue data feature.