- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
02-20-2023 11:55 PM
Hello to everyone!
I am trying to read delta table as a streaming source using spark. But my microbatches are disbalanced - one very small and the other are very huge. How I can limit this?
I used different configurations with maxBytesPerTrigger and maxFilesPerTrigger, but nothing changes, batch size is always the same.
Are there any ideas?
df = spark \
.readStream \
.format("delta") \
.load("...")
df \
.writeStream \
.outputMode("append") \
.option("checkpointLocation", "...") \
.table("...")
Kind Regards
- Labels:
-
Spark structured streaming
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
02-21-2023 04:12 AM
besides the parameters you mention, I don't know of any other which controls the batch size.
did you check if the delta table is not horribly skewed?
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
02-27-2023 08:52 AM
Thanks, you are right! Data was very skewed