Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
11-11-2025 08:44 AM
First of all, you are using append-only reads, which means that every time your stream triggers, Spark will process the entire Delta snapshot rather than just the changes.
That’s why you’re observing the memory usage increase after each run, it’s not a bug, it’s how Spark Structured Streaming works under the hood.
Use instead Delta Change Data Feed (CDF)
Change Data Feed (CDF) is a built-in Delta feature that lets you read only the changes (inserts, updates, deletes) instead of the full dataset.
When you enable it, Spark treats the Delta table as a true incremental stream source.
https://docs.databricks.com/aws/en/delta/delta-change-data-feed