- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
a month ago
I completely understand your concern, and I think this is a valid question.
If one ingestion pipeline is permanently associated with one ingestion gateway, then from an operational perspective, 20–30 source connections could indeed result in 20–30 ingestion pipelines. At first glance, that does sound like additional maintenance overhead, especially as the number of source systems continues to grow.
However, I think the key is to differentiate between having multiple pipeline instances and maintaining multiple independent pipeline implementations.
For example, I would try to manage this using a common YAML/template and metadata-driven configuration. The ingestion pattern can remain standard, while the gateway, connection, tables and target configuration are supplied per source.
Think of it as one reusable ingestion pattern, with each server or gateway having its own configuration and pipeline instance.
So you may still have 30 pipelines because of the Lakeflow Connect/gateway architecture, but ideally you should not have 30 different pipelines that need to be manually developed and maintained.
For example, the metadata/configuration could define:
Ingestion gateway
Source connection
Source database/schema
Tables to ingest
Target catalog/schema
Load configuration
A standard deployment process could then create and update the relevant pipeline configurations consistently.
I also think there are some advantages to having the pipelines separated by connection/gateway. For example, issues with one source connection are more isolated and do not necessarily impact ingestion from other source systems. It can also make monitoring, troubleshooting and access management more clearly scoped to a particular source.
That said, I agree that the number of pipeline resources can become an operational challenge at scale. In my view, this is where automation becomes important. I would avoid managing 30 pipelines manually through the UI and instead use a standardised configuration, source control and CI/CD approach to manage them.
So, the way I see it is, you may have multiple pipeline instances, but you should aim for a single reusable ingestion pattern and automate the creation and configuration of those pipelines.
That way, onboarding a new SQL Server becomes mainly a configuration/deployment exercise rather than creating and maintaining a completely new ingestion solution each time.