cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

Anyone migrated from legacy CDF to Auto CDF yet?

data_pulse
New Contributor III

Databricks mentions in the DBR 19 release notes that Auto CDF removes write-time overhead and can make MERGE and UPDATE operations about 15% faster on tables queried for changes

For anyone who has moved an existing pipeline from legacy CDF to Auto CDF:

1) What did your before/after numbers look like for MERGE or UPDATE duration? Did you get any improvement?
2) Have you seen any meaningful reduction in storage growth as well? I'm especially interested in tables with frequent MERGEs where legacy CDF was generating a noticeable amount of change data.
3) Hit with any practical migration issues with existing downstream streams/checkpoints?
4) Was the improvement large enough that you're actively moving all existing tables?

Would be great to hear where Auto CDF has actually made a noticeable difference in production and mostly in terms of: Merge Duration/ storage growth/ cost etc.

1 REPLY 1

ThomazNeto
Databricks Partner

Hi,

No production before/after numbers from me yet, we only have this on a couple of dev tables since GA on Sep 1, so I'll stick to what the docs say and what I'd measure. Someone with a month of prod data will hopefully chime in.

The main thing to understand is that there's nothing to "turn on". Auto CDF "computes row-level changes at query time using row tracking", so the only requirement on an existing Delta table is row tracking plus DBR 19. Migration is literally two statements:

ALTER TABLE my_table SET TBLPROPERTIES (delta.enableRowTracking = true);
ALTER TABLE my_table UNSET TBLPROPERTIES ('delta.enableChangeDataFeed');

Readers don't change: same table_changes() and readChangeFeed, batch and streaming.
link doc
link doc

On your questions:

1 and 2. The 15% and "reduces storage costs" are Databricks' claims; the mechanism supports them (no _change_data files written on MERGE/UPDATE anymore), so on tables with frequent MERGEs the storage side should be the visible one. To measure it yourself before and after: DESCRIBE HISTORY gives you operationMetrics.executionTimeMs per MERGE and the bytes added; system.query.history gives the wall time. Compare a week of each.

  1. This is where the docs are silent, and where I'd be careful. Two known costs: enabling row tracking on an existing table rewrites metadata for every row, the docs say it "might result in the creation of multiple new versions of the table and take a significant amount of time", and it must not run concurrently with writers (MetadataChangedException), so pause the streams for it. What the docs don't say is whether an existing readChangeFeed checkpoint continues cleanly over the switch, or what happens to already-materialized change history after UNSET. I'd test the migration on a copy with the real downstream stream attached before touching production. One useful bit: the docs say a gap in legacy CDF history "will not be queryable. Use automatic change data feed to query changes during the interval", which suggests Auto CDF can read across the switch on row-tracked tables, but that's the closest they get to a guarantee.
  2. Not yet, for two reasons from the limitations list: Auto CDF "isn't supported on tables with row filters or column masks" (so it's out for our ABAC-governed tables), and only Databricks readers can consume it, external Delta clients can't. The row tracking writer protocol upgrade also matters if anything outside Databricks writes to the table. Where none of that applies, I'd move MERGE-heavy tables first; the docs themselves say "Databricks recommends that you migrate to automatic change data feed".
    link doc

If you run the numbers, please post them here.

Thomaz A. Rossito Neto
Principal Data Architect & AI Strategy — CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto