Lakeflow connect

ram_11
New Contributor
Spoiler
 

What is extractor_sql-server-conn_cdc_sink in the Ingestion gateway pipeline in lakeflow connect and is it crreated by deafualta and how does it work?

K_Anudeep
Databricks Employee
Databricks Employee

Hi @ram_11 !

The extractor_sql-server-conn_cdc_sink is the shared CDC sink that the SQL Server CDC extractor job inside the IG pipeline writes to.

Example screenshot below:

K_Anudeep_1-1791011265117.png

 

 

What actually happens:

  1. The IG extractor (the job/process) reads change rows from each table’s capture instance / CT table.
  2. When the IG pipelines run, we see that the per-table CDC flows are defined 
  3. All of those flows append JSON into one shared sink: extractor_<connection-name>_cdc_sink.
  4. This sink is a UC volume cdc/ directory inside your staging location where your ingesting pipeline reads from later.

Also, when you start your IG pipleine for the first time, you should see extractor_sql-server-conn_snaphsot_sink as well, which is a sink defined, one for each table where the snapshot extractor writes the initial/full-load Parquet files.

I hope this answers your question.

Anudeep

View solution in original post

ram_11
New Contributor

Thanks for the answer..it's now clear..I also followed the doc and understood!