a week ago
Hi Team,
I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.
I created a SQL Server ingestion pipeline using:
The SQL connection, gateway definition, and pipeline configuration were all created successfully.
When Lakeflow attempts to start the CDC gateway, the compute provisioning fails with the following error:
Failed to add 2 workers to the compute. Will attempt retry: true. Reason:Cloud Provider Resource Stockout
The VM size you are specifying is not available. To reduce future stockout errors, enable flexible node types if not already enabled so Databricks can automatically fall back to compatible instance types when capacity is limited. Note: flexible node types are not available for GPU instance types. [details] SkuNotAvailable: The requested VM size for resource 'Following SKUs have failed for Capacity Restrictions: Standard_E4d_v4' is currently not available in location 'australiaeast'. Please try another size or deploy to a different location or different zone.
I inspected the generated compute configuration and found that Lakeflow is attempting to provision:
The issue appears to be that Azure Australia East currently cannot allocate Standard_E4d_v4.
Is there a supported way to configure or override the VM types used by a Lakeflow Connect SQL Server CDC Gateway?
Specifically:
Any guidance would be greatly appreciated.
a week ago
@Mado this looks like a gateway compute capacity issue rather than a CDC issue.
Databricks does support customizing SQL Server gateway compute through a custom Job Compute policy/API. It is a good practice to keep workers small; requires at least 8 cores on the driver for efficient extraction.
Also, "Standard_E4d_v4" currently has no compatible Flexible Node Type fallback, so enabling Flexible Node Types won't help with this particular SKU.
I'd use an available v5 driver with 8+ cores through a custom policy/API, plus a small worker SKU available in Astralia East, rather than relying on the UI-generated defaults.
a week ago - last edited a week ago
There's no UI-exposed setting in Lakeflow Connect to choose the gateway's VM SKU today — the driver/worker types (Standard_E4d_v4 / Standard_D4s_v4) appear to be fixed defaults for the standard CDC gateway, not a configurable field in the connection/gateway/pipeline creation flow. That matches what you found in your investigation.
On Flexible Node Types not helping here: worth noting that the standard CDC gateway runs on classic compute in your workspace VNet (per Databricks' architecture docs for managed connectors), not serverless. Flexible node types is a compute-policy-level feature, and it's plausible the gateway's internally-managed compute isn't actually inheriting your workspace-level flexible-node-types setting the way a normal job cluster would — which would explain why it's still stuck retrying the exact same unavailable SKU instead of falling back. That's consistent with what you're seeing rather than a contradiction of it.
A reported workaround (not officially documented, but seen in the wild): at least one similar Azure capacity/stockout case for a Lakeflow Connect gateway was resolved by pinning the gateway's driver/worker node types via a compute policy applied to the gateway's underlying pipeline compute, rather than relying on the default. If gateway pipelines expose a clusters spec (node_type_id / driver_node_type_id) via the Pipelines REST API or a Declarative Automation Bundle definition, setting that explicitly to a v5 family you know has capacity (e.g., Standard_E4ds_v5 / Standard_D4ds_v5) is worth trying — this wouldn't necessarily be exposed through the Lakeflow Connect UI wizard, so you may need to go through the API/DAB path to set it directly on the gateway pipeline resource.
A more structural alternative: Databricks has an Integrated CDC (Beta) architecture for SQL Server that collapses the two-pipeline (gateway + ingestion) model into a single pipeline running on serverless compute instead of a gateway with fixed VM sizing — this would sidestep the stockout problem entirely since there's no specific SKU to provision. Trade-offs while it's in Beta: ~30-minute max runtime per pipeline update (resumes on the next scheduled run if there's more to process), and it doesn't yet support row filtering or automated data-type schema evolution (column add/delete is supported). If your CDC volume and latency needs tolerate that, it's worth evaluating as a way around this specific issue rather than fighting the gateway's VM selection.
a week ago
Hi Mado,
Ivan is right on the diagnosis. Adding the concrete "how", since the docs only give it in fragments.
Two constraints to respect: the driver needs at least 8 cores ("The minimum requirement for the driver node is 8 cores"), so Standard_E4ds_v5 is too small for the driver, go E8ds_v5 or larger; and the gateway must be classic compute, serverless isn't an option for it.
link
Australia East stockouts on v4 E-series are a recurring thing, so pinning v5 through the policy is the right long-term fix, not a workaround.
Wednesday
Thanks @ThomazNeto for your help.
I created a policy based on your instructions.
Could you please advise how I can configure the Gateway Pipeline to use this policy?
If there are any additional settings or steps required to ensure the Gateway Pipeline uses the policy during execution, I'd appreciate your guidance.
Thanks again for your help.
Kind regards,
Mohammad
a week ago
Hi @Mado
One important distinction here is between the standard MySQL CDC architecture and the newer Integrated CDC pipeline.
For the Integrated CDC approach, Classic compute is expected — Serverless compute is currently not supported.
The current architecture is roughly:
MySQL
↓
Unity Catalog Connection
↓
Integrated CDC Pipeline
↓
Staging Volume
↓
Destination Streaming Tables
A few things to check:
Workspace feature enablement
Integrated CDC is currently a Beta feature and requires workspace-level enablement.
Compute
The Integrated CDC pipeline requires Classic compute. Setting serverless: false is expected; Serverless isn't supported for this pipeline type.
MySQL configuration
Make sure binary logging is enabled with:
binlog_format = ROW
binlog_row_image = FULL
MySQL replication user
The connection user needs the required replication privileges.
Pipeline creation
Integrated CDC currently supports API/CLI/notebook/DAB-based creation; UI-based pipeline creation isn't supported yet.
One useful point is that Integrated CDC is different from the standard gateway-based architecture. The standard approach uses a separate ingestion gateway, while Integrated CDC combines extraction and application into a single pipeline update.
So if your requirement is specifically MySQL Integrated CDC, having only Classic compute should not itself be the blocker — Classic compute is actually the required compute type for this feature.
If you're seeing a specific error while creating the pipeline, sharing the error message would help narrow down whether the issue is feature enablement, connection configuration, or source-side CDC setup.
References:
Databricks MySQL Integrated CDC documentation: https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/mysql-integrated-pipeline
MySQL source configuration: https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/mysql-source-setup
a week ago
Hi @Mado
As of now, Lakeflow Connect's CDC Gateway VM types (driver/worker SKUs) are not configurable at the pipeline or connection level — they are managed internally by the platform. If the default SKU (Standard_E4d_v4) is unavailable in your region due to stockout, your best options are:
(1) retry provisioning later as Azure capacity fluctuates,
(2) contact Databricks Support directly to request they investigate enabling v5 SKU fallback for Australia East gateways, since "Flexible node types" alone doesn't appear to be resolving this for gateway compute specifically. This is a known pain point with managed/serverless-style compute layers where regional stockouts can't be self-serviced around.
yesterday - last edited yesterday
You can pin the gateway's node types with that policy. The SQL Server ingestion page lists a custom gateway policy as "API only". Attach it in the gateway pipeline's clusters setting, with apply_policy_default_values set to true so the policy's defaults are applied when you use the Pipelines API (pipeline compute) :
"clusters": [
{
"label": "default",
"policy_id": "<your-policy-id>",
"apply_policy_default_values": true
}
]
For the gateway the wizard created, copy the contents of the spec field from databricks pipelines get <gateway-id> -o json into spec.json, replace its clusters field with the entry above while keeping every other field, and send it back with databricks pipelines update <gateway-id> --json @spec.json (Edit a pipeline).
A running continuous gateway restarts automatically with the updated configuration (configuration changes). If the gateway is not running after the edit, start it with databricks pipelines start-update <gateway-id> (restart the ingestion gateway). If it is still stuck starting after 15 minutes, the SQL Server FAQ says to open a support ticket.