<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Multiple gateway pipeline for same database in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168122#M55849</link>
    <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250813"&gt;@srikanthp24&lt;/a&gt;&amp;nbsp;,&lt;/P&gt;
&lt;P&gt;To your core question: multiple gateway pipelines against the same source database don't interfere with each other. Each gateway keeps its own independent state, its own staging volume, and its own CDC/change-tracking cursor, so a separate Dev, UAT, and Prod gateway pointing at the same SQL Server won't step on one another. As long as each pair uses distinct names, staging, and destination schemas (as noted above), your existing Dev pipeline is unaffected. Databricks even points at this pattern for resolving table-name conflicts: &lt;SPAN&gt;&lt;SPAN draggable="true"&gt;&lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server-troubleshoot" rel="noopener noreferrer" target="_blank"&gt;create multiple gateway-pipeline pairs writing to different destination schemas&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;.&lt;BR /&gt;&lt;BR /&gt;The one thing I'd watch that isn't obvious: the risk with several environments isn't on the Databricks side, it's the SQL Server CDC cleanup job. It runs on a retention schedule and isn't aware of individual consumers, so your least-active environment sets the floor. If a UAT or Dev gateway sits stopped past the CDC retention window, its cursor ends up pointing at change data that's already been purged, which forces a full refresh of the affected tables. That's why the &lt;SPAN&gt;&lt;SPAN draggable="true"&gt;&lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server/limitations" rel="noopener noreferrer" target="_blank"&gt;gateway must run continuously&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;, and why you should size source-side CDC retention for the slowest gateway rather than the typical one.&lt;BR /&gt;&lt;BR /&gt;Two smaller notes. Adding more gateways doesn't consume extra CDC capture instances, since they all read the same capture instance rather than creating their own, so you won't hit SQL Server's two-instance-per-table limit just by adding environments. And since each environment adds a continuous reader against the source, &lt;SPAN&gt;&lt;SPAN draggable="true"&gt;&lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server-source-setup" rel="noopener noreferrer" target="_blank"&gt;change tracking is recommended over CDC for any table with a primary key&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt; to keep source load down. If you're greenfielding UAT/Prod, the newer Integrated CDC (Beta) mode also collapses the gateway and pipeline into one, worth a look.&lt;/P&gt;</description>
    <pubDate>Wed, 09 Sep 2026 18:53:02 GMT</pubDate>
    <dc:creator>stbjelcevic</dc:creator>
    <dc:date>2026-09-09T18:53:02Z</dc:date>
    <item>
      <title>Multiple gateway pipeline for same database</title>
      <link>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168103#M55838</link>
      <description>&lt;P&gt;Hi Team,&lt;/P&gt;&lt;P&gt;To do the POC on Lakeflow Connect Ingestion Pipeline, we directly connected the sql server prod DB and build the pipeline it is running successfully. Now, My question is we need to move that to UAT and Prod. In Dev we have created througn UI Wizard. In UAT I am planning to creating using the DABs. Is that impact existing dev pipeline or Is okay to create the multiple gateway pipeline in different&amp;nbsp; environment pointing to same source database.&lt;/P&gt;&lt;P&gt;Can anyone guide me on this clearly.&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 09 Sep 2026 14:37:27 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168103#M55838</guid>
      <dc:creator>srikanthp24</dc:creator>
      <dc:date>2026-09-09T14:37:27Z</dc:date>
    </item>
    <item>
      <title>Re: Multiple gateway pipeline for same database</title>
      <link>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168107#M55839</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250813"&gt;@srikanthp24&lt;/a&gt;&amp;nbsp;&lt;BR /&gt;Before Moving, Please consider the following Questions&lt;BR /&gt;&lt;BR /&gt;Can the source SQL Server handle the additional CDC read load?&amp;nbsp;(While adding a second pipeline will not break the Dev pipeline, it doubles the CDC query load on your production SQL Server.)&lt;BR /&gt;What is the CDC log retention window on SQL Server?&amp;nbsp;Ensure your SQL Server CDC cleanup retention period is set long enough (e.g., 3–7 days). If one environment (e.g., UAT) is paused for several days and CDC logs are cleaned up on SQL Server before UAT resumes, the UAT pipeline will throw an out-of-range offset error and require a re-sync.&lt;BR /&gt;Are service accounts and network paths validated?&lt;BR /&gt;&lt;BR /&gt;My Idea is&amp;nbsp;Deploying via DABs (Best Practice Workflow)&lt;BR /&gt;&lt;BR /&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 09 Sep 2026 15:15:10 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168107#M55839</guid>
      <dc:creator>Satyasai</dc:creator>
      <dc:date>2026-09-09T15:15:10Z</dc:date>
    </item>
    <item>
      <title>Re: Multiple gateway pipeline for same database</title>
      <link>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168108#M55840</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250813"&gt;@srikanthp24&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Yes, creating UAT/PROD with DAB is a good approach and should not impact the existing DEV pipeline, as long as each environment is deployed as a separate resource with its own names, staging location and destination schema/catalog.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Typical Example:&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;DEV -&amp;gt; gateway-dev -&amp;gt; ingestion-dev -&amp;gt; dev_catalog&lt;BR /&gt;UAT -&amp;gt; gateway-uat -&amp;gt; ingestion-uat -&amp;gt; uat_catalog&lt;BR /&gt;PROD -&amp;gt; gateway-prod -&amp;gt; ingestion-prod -&amp;gt; prod_catalog&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;One important limitation&lt;/STRONG&gt;: each ingestion pipeline must have exactly one gateway, and a gateway cannot be shared across multiple ingestion pipelines. Can see more limitations &lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server-limits" target="_self"&gt;here&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;BR /&gt;If DEV/UAT/PROD all point to the same SQL Server source, it works, but remember that each gateway is a separate CDC reader, so should consider the additional load on the source and keep staging/checkpoint state separate.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;On question of Creating Pipeline via DAB in Dev:&lt;/STRONG&gt;&amp;nbsp;The pipeline created with DAB can coexist with the existing UI created DEV pipeline as long as the bundle uses a &lt;STRONG&gt;different gateway name and ingestion pipeline name&lt;/STRONG&gt;. Databricks requires unique names when creating both the ingestion pipeline and the gateway.&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="data_pulse_0-1788966666486.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30933iD1BA03B3E72EA801/image-size/medium?v=v2&amp;amp;px=400" role="button" title="data_pulse_0-1788966666486.png" alt="data_pulse_0-1788966666486.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Alternatively,&lt;/STRONG&gt;&amp;nbsp;For the existing DEV pipeline created through the UI, if you later want to manage that same pipeline through DABs, use bundle bundle deployment bind rather than creating a new bundle resource.&lt;/P&gt;&lt;P&gt;Eg: If existing Dev pipeline Id is : &lt;STRONG&gt;123abc&amp;nbsp;&lt;/STRONG&gt;and bundle resource looks like&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;resources:
  pipelines:
    dev_ingestion:
      name: sqlserver-dev-ingestion&lt;/LI-CODE&gt;&lt;P&gt;Then bind that resource to the existing DEV pipeline:&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;databricks bundle deployment bind
  dev_ingestion
  123abc
  --target dev&lt;/LI-CODE&gt;&lt;P&gt;Same for prod too, Can bind the Existing Pipelines (As above).&lt;/P&gt;</description>
      <pubDate>Wed, 09 Sep 2026 15:19:27 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168108#M55840</guid>
      <dc:creator>data_pulse</dc:creator>
      <dc:date>2026-09-09T15:19:27Z</dc:date>
    </item>
    <item>
      <title>Re: Multiple gateway pipeline for same database</title>
      <link>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168122#M55849</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/250813"&gt;@srikanthp24&lt;/a&gt;&amp;nbsp;,&lt;/P&gt;
&lt;P&gt;To your core question: multiple gateway pipelines against the same source database don't interfere with each other. Each gateway keeps its own independent state, its own staging volume, and its own CDC/change-tracking cursor, so a separate Dev, UAT, and Prod gateway pointing at the same SQL Server won't step on one another. As long as each pair uses distinct names, staging, and destination schemas (as noted above), your existing Dev pipeline is unaffected. Databricks even points at this pattern for resolving table-name conflicts: &lt;SPAN&gt;&lt;SPAN draggable="true"&gt;&lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server-troubleshoot" rel="noopener noreferrer" target="_blank"&gt;create multiple gateway-pipeline pairs writing to different destination schemas&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;.&lt;BR /&gt;&lt;BR /&gt;The one thing I'd watch that isn't obvious: the risk with several environments isn't on the Databricks side, it's the SQL Server CDC cleanup job. It runs on a retention schedule and isn't aware of individual consumers, so your least-active environment sets the floor. If a UAT or Dev gateway sits stopped past the CDC retention window, its cursor ends up pointing at change data that's already been purged, which forces a full refresh of the affected tables. That's why the &lt;SPAN&gt;&lt;SPAN draggable="true"&gt;&lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server/limitations" rel="noopener noreferrer" target="_blank"&gt;gateway must run continuously&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;, and why you should size source-side CDC retention for the slowest gateway rather than the typical one.&lt;BR /&gt;&lt;BR /&gt;Two smaller notes. Adding more gateways doesn't consume extra CDC capture instances, since they all read the same capture instance rather than creating their own, so you won't hit SQL Server's two-instance-per-table limit just by adding environments. And since each environment adds a continuous reader against the source, &lt;SPAN&gt;&lt;SPAN draggable="true"&gt;&lt;A href="https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/sql-server-source-setup" rel="noopener noreferrer" target="_blank"&gt;change tracking is recommended over CDC for any table with a primary key&lt;/A&gt;&lt;/SPAN&gt;&lt;/SPAN&gt; to keep source load down. If you're greenfielding UAT/Prod, the newer Integrated CDC (Beta) mode also collapses the gateway and pipeline into one, worth a look.&lt;/P&gt;</description>
      <pubDate>Wed, 09 Sep 2026 18:53:02 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168122#M55849</guid>
      <dc:creator>stbjelcevic</dc:creator>
      <dc:date>2026-09-09T18:53:02Z</dc:date>
    </item>
    <item>
      <title>Re: Multiple gateway pipeline for same database</title>
      <link>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168172#M55856</link>
      <description>&lt;P&gt;Thanks for the clarification&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/167034"&gt;@stbjelcevic&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The docs say the SQL Server gateway must run continuously and can resume from its previous point as long as the required source logs still exist.&amp;nbsp;Is there any documented way to determine how far behind a gateway can safely fall before a full refresh is required or is that entirely governed by the SQL Server CDC/change-tracking retention configuration?&lt;/P&gt;&lt;P&gt;Also is there a recommended metric/event/log to monitor the gateway’s current CDC position/ lag against the old retained source log, so we can alert before the retention window is exceeded?&lt;/P&gt;</description>
      <pubDate>Thu, 10 Sep 2026 06:55:47 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/multiple-gateway-pipeline-for-same-database/m-p/168172#M55856</guid>
      <dc:creator>data_pulse</dc:creator>
      <dc:date>2026-09-10T06:55:47Z</dc:date>
    </item>
  </channel>
</rss>

