cancel
Showing results for 
Search instead for 
Did you mean: 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results for 
Search instead for 
Did you mean: 

How can I configure Lakeflow Connect SQL Server CDC Gateway to use a desired VM type?

Mado
Valued Contributor II

Hi Team,

I'm evaluating Databricks Lakeflow Connect for SQL Server CDC ingestion and have run into a gateway provisioning issue.

Environment

  • Region: Australia East
  • Source: Azure SQL Database
  • CDC enabled successfully on the database and source table
  • SQL Server connection created successfully in Lakeflow Connect

     

What I configured

I created a SQL Server ingestion pipeline using:

  • Change Data Capture (CDC)
  • Ingestion pipeline
  • Ingestion gateway

    The SQL connection, gateway definition, and pipeline configuration were all created successfully.

     

 

Problem

When Lakeflow attempts to start the CDC gateway, the compute provisioning fails with the following error:

 

Message

Failed to add 2 workers to the compute. Will attempt retry: true. Reason:Cloud Provider Resource Stockout

Help

The VM size you are specifying is not available. To reduce future stockout errors, enable flexible node types if not already enabled so Databricks can automatically fall back to compatible instance types when capacity is limited. Note: flexible node types are not available for GPU instance types. [details] SkuNotAvailable: The requested VM size for resource 'Following SKUs have failed for Capacity Restrictions: Standard_E4d_v4' is currently not available in location 'australiaeast'. Please try another size or deploy to a different location or different zone. 

 
 
Investigation

I inspected the generated compute configuration and found that Lakeflow is attempting to provision:

 
Driver: Standard_E4d_v4
 
Worker: Standard_D4s_v4
 
I also verified the following:
  • Azure SQL connectivity works.
  • CDC is functioning correctly.
  • Flexible Node Types has been enabled on the workspace.

    The issue appears to be that Azure Australia East currently cannot allocate Standard_E4d_v4.

     

 

Question

Is there a supported way to configure or override the VM types used by a Lakeflow Connect SQL Server CDC Gateway?

Specifically:

  1. Can the gateway be configured to use v5 VM families (for example Standard_E4ds_v5 and Standard_D4ds_v5) instead of the default v4 families?
  2. Can this be configured during pipeline creation?
  3. If not, what is the recommended approach when the default gateway VM SKU is unavailable in the selected Azure region?

    Any guidance would be greatly appreciated.

    Mado_0-1790214214728.png

7 REPLIES 7

ivanvyd
New Contributor III

@Mado this looks like a gateway compute capacity issue rather than a CDC issue.

Databricks does support customizing SQL Server gateway compute through a custom Job Compute policy/API. It is a good practice to keep workers small; requires at least 8 cores on the driver for efficient extraction.

Also, "Standard_E4d_v4" currently has no compatible Flexible Node Type fallback, so enabling Flexible Node Types won't help with this particular SKU.

I'd use an available v5 driver with 8+ cores through a custom policy/API, plus a small worker SKU available in Astralia East, rather than relying on the UI-generated defaults.

Ivan Vydrin
Lead Software & AI Engineer · Tech Fabric LLC

aayush_410
New Contributor II

There's no UI-exposed setting in Lakeflow Connect to choose the gateway's VM SKU today — the driver/worker types (Standard_E4d_v4 / Standard_D4s_v4) appear to be fixed defaults for the standard CDC gateway, not a configurable field in the connection/gateway/pipeline creation flow. That matches what you found in your investigation.

On Flexible Node Types not helping here: worth noting that the standard CDC gateway runs on classic compute in your workspace VNet (per Databricks' architecture docs for managed connectors), not serverless. Flexible node types is a compute-policy-level feature, and it's plausible the gateway's internally-managed compute isn't actually inheriting your workspace-level flexible-node-types setting the way a normal job cluster would — which would explain why it's still stuck retrying the exact same unavailable SKU instead of falling back. That's consistent with what you're seeing rather than a contradiction of it.

A reported workaround (not officially documented, but seen in the wild): at least one similar Azure capacity/stockout case for a Lakeflow Connect gateway was resolved by pinning the gateway's driver/worker node types via a compute policy applied to the gateway's underlying pipeline compute, rather than relying on the default. If gateway pipelines expose a clusters spec (node_type_id / driver_node_type_id) via the Pipelines REST API or a Declarative Automation Bundle definition, setting that explicitly to a v5 family you know has capacity (e.g., Standard_E4ds_v5 / Standard_D4ds_v5) is worth trying — this wouldn't necessarily be exposed through the Lakeflow Connect UI wizard, so you may need to go through the API/DAB path to set it directly on the gateway pipeline resource.

A more structural alternative: Databricks has an Integrated CDC (Beta) architecture for SQL Server that collapses the two-pipeline (gateway + ingestion) model into a single pipeline running on serverless compute instead of a gateway with fixed VM sizing — this would sidestep the stockout problem entirely since there's no specific SKU to provision. Trade-offs while it's in Beta: ~30-minute max runtime per pipeline update (resumes on the next scheduled run if there's more to process), and it doesn't yet support row filtering or automated data-type schema evolution (column add/delete is supported). If your CDC volume and latency needs tolerate that, it's worth evaluating as a way around this specific issue rather than fighting the gateway's VM selection.

Aayush Sharma

ThomazNeto
Databricks Partner

Hi Mado,

Ivan is right on the diagnosis. Adding the concrete "how", since the docs only give it in fragments.

  1. Can the gateway use other VM families? Yes, but not from the wizard. The UI walkthrough has no compute step for the gateway; the docs list the policy route as "Unrestricted permissions to create clusters, or a custom policy (API only)". Two ways to set it:
  • Cluster policy: family Job Compute, with the overrides the docs require ("cluster_type": fixed "dlt", "runtime_engine": fixed "STANDARD", "num_workers": unlimited, default 1) plus your node types, e.g. "driver_node_type_id": fixed "Standard_E8ds_v5" and "node_type_id": fixed "Standard_D4ds_v5". The docs' own Azure example is Standard_E64d_v4 driver and Standard_F4s workers, and they say workers "do not impact gateway performance", so keep them tiny.
  • Directly in the gateway pipeline definition (bundles or Pipelines API): the gateway is a pipeline object and the docs show a commented clusters block on it, with label "default", autoscale, node_type_id and policy_id. That's where you pin the SKU without a policy.
    link

Two constraints to respect: the driver needs at least 8 cores ("The minimum requirement for the driver node is 8 cores"), so Standard_E4ds_v5 is too small for the driver, go E8ds_v5 or larger; and the gateway must be classic compute, serverless isn't an option for it.
link

  1. During creation? Only via API, bundles or the Terraform examples repo linked from the docs. Editing the clusters block afterwards isn't documented for the gateway specifically, so I'd recreate the gateway from YAML for the POC rather than patch the UI-generated one.
  2. Why Flexible Node Types didn't save you: Standard_E4d_v4 (the non-"s" variant) simply isn't in the compatibility table. Standard_E4ds_v4 falls back to E4ds_v5 and L4s, Standard_E4d_v4 has no entry, so there's nothing to fall back to. Pick the "ds" v5 SKUs, which are in the table and have their own fallbacks.
    link

Australia East stockouts on v4 E-series are a recurring thing, so pinning v5 through the policy is the right long-term fix, not a workaround.

Thomaz A. Rossito Neto
Principal Data Architect & AI Strategy — CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto

Mado
Valued Contributor II

Thanks @ThomazNeto  for your help. 

I created a policy based on your instructions.

Could you please advise how I can configure the Gateway Pipeline to use this policy?

If there are any additional settings or steps required to ensure the Gateway Pipeline uses the policy during execution, I'd appreciate your guidance.

Thanks again for your help.

Kind regards,
Mohammad

mukul1409
Contributor II

Hi @Mado 

One important distinction here is between the standard MySQL CDC architecture and the newer Integrated CDC pipeline.

For the Integrated CDC approach, Classic compute is expected — Serverless compute is currently not supported.

The current architecture is roughly:

MySQL
↓
Unity Catalog Connection
↓
Integrated CDC Pipeline
↓
Staging Volume
↓
Destination Streaming Tables

A few things to check:

Workspace feature enablement
Integrated CDC is currently a Beta feature and requires workspace-level enablement.

Compute
The Integrated CDC pipeline requires Classic compute. Setting serverless: false is expected; Serverless isn't supported for this pipeline type.

MySQL configuration
Make sure binary logging is enabled with:

binlog_format = ROW

binlog_row_image = FULL

MySQL replication user
The connection user needs the required replication privileges.

Pipeline creation
Integrated CDC currently supports API/CLI/notebook/DAB-based creation; UI-based pipeline creation isn't supported yet.

One useful point is that Integrated CDC is different from the standard gateway-based architecture. The standard approach uses a separate ingestion gateway, while Integrated CDC combines extraction and application into a single pipeline update.

So if your requirement is specifically MySQL Integrated CDC, having only Classic compute should not itself be the blocker — Classic compute is actually the required compute type for this feature.

If you're seeing a specific error while creating the pipeline, sharing the error message would help narrow down whether the issue is feature enablement, connection configuration, or source-side CDC setup.

References:
Databricks MySQL Integrated CDC documentation: https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/mysql-integrated-pipeline
MySQL source configuration: https://docs.databricks.com/aws/en/ingestion/lakeflow-connect/mysql-source-setup

Mukul Chauhan

Khasim_1
New Contributor III

Hi @Mado 

As of now, Lakeflow Connect's CDC Gateway VM types (driver/worker SKUs) are not configurable at the pipeline or connection level — they are managed internally by the platform. If the default SKU (Standard_E4d_v4) is unavailable in your region due to stockout, your best options are:

(1) retry provisioning later as Azure capacity fluctuates,

(2) contact Databricks Support directly to request they investigate enabling v5 SKU fallback for Australia East gateways, since "Flexible node types" alone doesn't appear to be resolving this for gateway compute specifically. This is a known pain point with managed/serverless-style compute layers where regional stockouts can't be self-serviced around.

Data Architect | 13 Years Domain Expertise | Databricks SA Champion Cohort

AbhilashNagilla
Databricks Employee
Databricks Employee

You can pin the gateway's node types with that policy. The SQL Server ingestion page lists a custom gateway policy as "API only". Attach it in the gateway pipeline's clusters setting, with apply_policy_default_values set to true so the policy's defaults are applied when you use the Pipelines API (pipeline compute) : 

"clusters": [
  {
    "label": "default",
    "policy_id": "<your-policy-id>",
    "apply_policy_default_values": true
  }
]

For the gateway the wizard created, copy the contents of the spec field from databricks pipelines get <gateway-id> -o json into spec.json, replace its clusters field with the entry above while keeping every other field, and send it back with databricks pipelines update <gateway-id> --json @spec.json (Edit a pipeline).

A running continuous gateway restarts automatically with the updated configuration (configuration changes). If the gateway is not running after the edit, start it with databricks pipelines start-update <gateway-id> (restart the ingestion gateway). If it is still stuck starting after 15 minutes, the SQL Server FAQ says to open a support ticket.