Widening the NSG service tag to all regions is a good first step but it rarely fixes this on its own. The AzureDatabricks service tag only covers the control plane traffic. What most people miss is that during bootstrap, the VM also needs to reach Azure Blob Storage endpoints to pull down the Databricks runtime artifacts, and those are not included in that service tag at all. That is most likely why widening the tag made no difference.
Try going through these one by one:
1. Quickly rule out a regional outage
Check https://status.azuredatabricks.net/ before going too deep into troubleshooting. Southeast Asia has had occasional blips and it is worth a 30-second check.
2. Allowlist the blob storage FQDNs โ this is usually the fix
You need outbound HTTPS access to region-specific storage endpoints like dbartifactsprodseap.blob.core.windows.net.
The full list for your region is here:
https://learn.microsoft.com/en-us/azure/databricks/resources/supported-regions
3. Double check your VNet injection setup
If you are on a customer-managed VNet, make sure:
- Both subnets are delegated to Microsoft.Databricks/workspaces
- The Databricks-managed NSG rules are still intact and have not been accidentally removed
- There is a valid outbound internet path - NAT gateway, UDR to a firewall, or default route
4. Test DNS from within the same subnet
Spin up a small test VM in the Databricks subnet and run nslookup against the control plane FQDN and the blob storage endpoints. If you have a custom DNS server in the VNet, it might be timing out on resolution and quietly eating into that 700-second bootstrap window before the VM even starts downloading anything.
5. Check your firewall or NVA if you have one
If traffic is routing through an Azure Firewall or any third-party NVA with TLS inspection enabled on storage traffic, the added latency can easily push bootstrap over the timeout. Try exempting the Databricks blob storage FQDNs from deep packet inspection and see if that changes anything.
6. Test connectivity directly from the subnet
From that test VM, run:
nslookup <control-plane-fqdn>
curl -v https://<artifact-storage-fqdn>
curl -v https://<workspace-url>
7. If everything checks out and it still fails
You simply raise a support ticket with Databricks and share the cluster ID, workspace ID, and exact timestamps. They can pull the backend bootstrap logs and tell you exactly which step timed out.
Also worth noting - the OnDemand in your error can sometimes indicate Azure capacity constraints in the region causing the VM to start slowly, which eats into the bootstrap window before your network config even comes into play.
Hope this helps!