- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
3 weeks ago
Nice writeup — this distinction between logical and physical deletion is exactly the kind of thing that trips people up when a "right to erasure" request lands on your desk and legal wants a straight yes/no answer.
Worth flagging that all of this soft-delete behavior only kicks in if deletion vectors are actually enabled on the table (delta.enableDeletionVectors = true, which is the default on newer runtimes/UC-managed tables now). Without DVs, DELETE just rewrites the Parquet files on the spot, so REORG...APPLY(PURGE) has nothing to do.
The retention window is the part I'd push back on a little — delta.deletedFileRetentionDuration isn't really a performance knob, it's your compliance gate. You can force it down with VACUUM ... RETAIN 0 HOURS and disabling the retention check, but that kills time travel and can break concurrent readers, so it's not something to just script — someone needs to actually sign off on giving up that safety in exchange for erasure speed.
Also worth turning on spark.databricks.delta.vacuum.logging.enabled before you need it — by default VACUUM doesn't log which files it deleted, and that log is exactly what you'd want to hand a DPO as proof the data is physically gone.
And honestly the part that breaks most erasure processes isn't any of this — it's everything downstream that still has a copy: Delta Sharing recipients, extracts, BI caches, CDC feeds, Change Data Feed itself. All of those need their own purge/retention story or the "deleted" record just keeps existing somewhere else.
So yes, in practice it ends up being three separate SLAs: logical removal is immediate, physical removal is bounded by retention + VACUUM, and downstream removal is tracked per system — and that last one is usually the slowest and riskiest leg.