Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
07-22-2025 09:01 AM - edited 07-22-2025 09:01 AM
You're facing a common issue with Spark's bad records handling.
Read CSV in PERMISSIVE mode and capture corrupt rows.df = spark.read
.option("mode", "PERMISSIVE")
.option("columnNameOfCorruptRecord", "_corrupt_record")
.format("csv")
.load("s3://your-bucket/path/")
later you can filter good and bad records from df.
LR