cross-region DR in Azure Databricks (24h RPO/RTO)
Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
3 weeks ago
Hi,
We're designing a DR strategy for an Azure Databricks platform and would appreciate guidance on current best practices for achieving approximately 24-hour RPO and RTO across Azure regions.
Our platform includes Unity Catalog, DAB, Jobs, Notebooks, SQL Warehouses, Lakeflow Declarative Pipelines, Materialized Views, Streaming Tables, Delta Sharing, and related security/governance components.
Our current approach is:
- Primary workspace in North Europe
- Standby workspace in West Europe
- Separate metastores per region
- DAB deployment to both workspaces
- DR resources kept paused until failover
- Nightly replication of business-critical data
- Recovery through a combination of replication, replay, and recomputation
A few questions:
- Is a warm-standby workspace the recommended pattern today?
- What is the preferred approach for Unity Catalog and metadata across regions?
- How are customers handling Lakeflow Pipelines, Streaming Tables, and Materialized Views given the lack of native cross-region replication?
- Is replay/recomputation the recommended recovery model for pipeline outputs?
- What are common approaches for SQL Warehouses, dashboards, alerts, and region-specific IDs?
- Are there reference architectures or proven customer patterns for 24h RPO/RTO on Azure Databricks?
- What DR limitations should we account for regarding Managed Volumes, MLflow models, Delta Shares, Vector Search, Genie assets, and similar services?
We're interested in both official Databricks guidance and real-world implementations from other customers.
Thanks in advance.
Labels:
- Labels:
-
Delta Lake
-
Spark