Options
- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
Tuesday
@Srini_Pesala You can deploy one workspace per region per cloud provider. Unity Catalog enforces certain limits on metastore per region. Data residency is handled natively through Databricks Geos and it ensures customer data is strictly processed and stored within the same geographic boundary as the workspace with no data leaving that boundary.
You can rely on Databricks-to-Databricks Open Sharing to handle data access across your AWS and Azure environments. It allows any workspace to seamlessly read shared tables from another workspace, generally regardless of the underlying cloud or region. You can physically replicate specific tables if they are heavily queried cross-region where cloud egress costs would become tricky to manage. You can handle it by setting up Delta Deep Clone pipelines that you can schedule using Lake flow Jobs at the required 4 hour cadence to push data to each region's storage.
DR strategy need to be split based on whether you are failing over within the same cloud or across clouds. For same-cloud, cross-region pairs like AWS-R1 to AWS-R2 or Azure-R1 to Azure-R2, Managed DR is the best. Databricks generally handles the replication of your UC metadata, managed table data and optionally workspace assets. It provides a stable URL allowing you to trigger the failover from the account console quickly.
You need to implement a new approach relying on Terraform or DAB for infrastructure-as-code, Delta Deep Clone for the underlying data replication and CI/CD pipelines to deploy code to both clouds simultaneously for cross cloud DR between AWS and Azure.
Use serverless compute generally if feasible. Unify identity management across the landscape by configuring account level SSO and SCIM synchronization. Managed the full architecture and controls via a central platform team using a Git-based infrastructure-as-code.