cancel
Showing results forย 
Search instead forย 
Did you mean:ย 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results forย 
Search instead forย 
Did you mean:ย 

Lakeflow connect

ram_11
New Contributor
Spoiler
 

What is extractor_sql-server-conn_cdc_sink in the Ingestion gateway pipeline in lakeflow connect and is it crreated by deafualta and how does it work?

1 REPLY 1

K_Anudeep
Databricks Employee
Databricks Employee

Hi @ram_11 !

The extractor_sql-server-conn_cdc_sink is the shared CDC sink that the SQL Server CDC extractor job inside the IG pipeline writes to.

Example screenshot below:

K_Anudeep_1-1791011265117.png

 

 

What actually happens:

  1. The IG extractor (the job/process) reads change rows from each tableโ€™s capture instance / CT table.
  2. When the IG pipelines run, we see that the per-table CDC flows are defined 
  3. All of those flows append JSON into one shared sink: extractor_<connection-name>_cdc_sink.
  4. This sink is a UC volume cdc/ directory inside your staging location where your ingesting pipeline reads from later.

Also, when you start your IG pipleine for the first time, you should see extractor_sql-server-conn_snaphsot_sink as well, which is a sink defined, one for each table where the snapshot extractor writes the initial/full-load Parquet files.

I hope this answers your question.

Anudeep