cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Streaming read fails with "Detected a data update" after a restatement job touches the source

Islam_hoti
New Contributor III

Hi everyone,

Looking for the right pattern here rather than a workaround.

Setup. DBR 15.4 LTS, Unity Catalog. A silver streaming table reads from a bronze Delta table with a normal streaming read. Bronze is append only in the ordinary course of business, so this has run cleanly for months.

The problem started when we introduced a monthly restatement job. Once a month, a correction process runs an UPDATE against bronze to fix a mis-mapped source code across historical rows. The next time the silver stream runs, it fails with the error about a data update being detected in the source table, and refuses to continue.

Compaction is not the issue. Regular OPTIMIZE runs have never broken this stream, which matches my understanding that compaction commits are marked as not changing data. It is specifically the restatement UPDATE that breaks it.

What I am weighing.

skipChangeCommits would let the stream continue, but it ignores the restated rows entirely, so silver would keep the wrong values forever. That seems worse than failing loudly.

Reading the change data feed instead of the table would let me propagate the corrections properly, but it changes the shape of the downstream logic, and I would need to handle update preimage and postimage records rather than plain appends.

Full refresh of silver after each restatement is the brute force option. It works, but silver is large enough that this is not something I want to do monthly.

Questions.

Is switching to a change data feed read the accepted answer here, or is there a pattern people prefer when corrections are rare and batched?

If I move to a CDF read, is there a clean way to cut over without reprocessing all of history, given silver is already correct up to a known point?

Does anyone restate through a separate correction stream instead, leaving the main append path untouched? That feels cleaner architecturally but I have not seen it written down anywhere.

Thanks.

1 ACCEPTED SOLUTION

Accepted Solutions

Khasim_1
New Contributor III

Hi @Islam_hoti ,

You have hit the classic 'Streaming vs. Mutation' wall. readStream on a Delta table assumes an append-only sequence; once an UPDATE happens, the commit versioning sequence is invalidated, hence the loud failure.

Here is how we reasoned through this for a similar production pipeline:

  1. Is CDF the accepted answer? Yes, for complex stateful logic, Change Data Feed (CDF) is the industry-standard way to handle updates in a stream. While it changes your downstream logic (needing to handle update_preimage and postimage), it is architecturally superior to a full refresh. It treats the 'Restatement' as a first-class event, which is exactly what it is.
  2. The Cutover Pattern: To move to CDF without re-processing history:
  • The 'Stateful' Cutover: You can use a startingVersion in your readStream to pick up the CDF from the exact commit version where you enabled it.
  • The Bridge: If you have 'Silver' data already correct up to Version X, you enable CDF on Bronze, point your new streaming read to the version after the restatement, and your Silver merge logic will naturally pick up the postimage and apply the correction.
  1. The 'Correction Stream' Pattern: You asked if anyone restates through a separate stream—we do, and it is the cleanest path for rare/batched updates.
  • Instead of UPDATEing the Bronze table (which is inherently anti-streaming), we treat the correction as a 'Correction Event'.
  • We write the correction to a separate Bronze_Corrections table.
  • The Silver stream uses a union (or coalesce) of the Bronze_Main and Bronze_Corrections.
  • The Silver logic (using a MERGE or Upsert) then processes these events naturally.

Architectural Recommendation: The reason your restatement job feels like a hack is because UPDATEing a source table is a batch-oriented operation being forced onto a streaming consumer. If you want to keep the architecture 'clean,' stop doing UPDATEs on the Bronze layer. Switch to an event-based correction stream. It preserves the immutability of your history, prevents the streaming pipeline from failing, and keeps your audit logs perfectly intact.

Data Architect | 13 Years Domain Expertise | Databricks SA Champion Cohort

View solution in original post

1 REPLY 1

Khasim_1
New Contributor III

Hi @Islam_hoti ,

You have hit the classic 'Streaming vs. Mutation' wall. readStream on a Delta table assumes an append-only sequence; once an UPDATE happens, the commit versioning sequence is invalidated, hence the loud failure.

Here is how we reasoned through this for a similar production pipeline:

  1. Is CDF the accepted answer? Yes, for complex stateful logic, Change Data Feed (CDF) is the industry-standard way to handle updates in a stream. While it changes your downstream logic (needing to handle update_preimage and postimage), it is architecturally superior to a full refresh. It treats the 'Restatement' as a first-class event, which is exactly what it is.
  2. The Cutover Pattern: To move to CDF without re-processing history:
  • The 'Stateful' Cutover: You can use a startingVersion in your readStream to pick up the CDF from the exact commit version where you enabled it.
  • The Bridge: If you have 'Silver' data already correct up to Version X, you enable CDF on Bronze, point your new streaming read to the version after the restatement, and your Silver merge logic will naturally pick up the postimage and apply the correction.
  1. The 'Correction Stream' Pattern: You asked if anyone restates through a separate stream—we do, and it is the cleanest path for rare/batched updates.
  • Instead of UPDATEing the Bronze table (which is inherently anti-streaming), we treat the correction as a 'Correction Event'.
  • We write the correction to a separate Bronze_Corrections table.
  • The Silver stream uses a union (or coalesce) of the Bronze_Main and Bronze_Corrections.
  • The Silver logic (using a MERGE or Upsert) then processes these events naturally.

Architectural Recommendation: The reason your restatement job feels like a hack is because UPDATEing a source table is a batch-oriented operation being forced onto a streaming consumer. If you want to keep the architecture 'clean,' stop doing UPDATEs on the Bronze layer. Switch to an event-based correction stream. It preserves the immutability of your history, prevents the streaming pipeline from failing, and keeps your audit logs perfectly intact.

Data Architect | 13 Years Domain Expertise | Databricks SA Champion Cohort