cancel
Showing results for 
Search instead for 
Did you mean: 
Administration & Architecture
Explore discussions on Databricks administration, deployment strategies, and architectural best practices. Connect with administrators and architects to optimize your Databricks environment for performance, scalability, and security.
cancel
Showing results for 
Search instead for 
Did you mean: 

Managed Disaster Recovery

yxnz1e
New Contributor

Hi, 

I have been doing testing on Managed DR, just few questions that i'm unable to find solution for:
1. I have observed that it takes more time in replication for empty catalogs rather than catalogs that have some data, does replication get stuck somewhere?
2. I had a metastore created for primary workspace which i eventually deleted and created new one, although the workspace picked up the new metastore, when i created failover group & tried to delete it, i got the error:

yxnz1e_0-1790318716914.png

apparently this metastore ID belongs to the metastore i deleted. Then if i delete the workspace will the failover group be deleted too?



2 REPLIES 2

ThomazNeto
Databricks Partner

Hi,

I've tested managed DR as well, and the docs cover less of this than you'd hope, so here's what they do say.

1. Slow replication on empty catalogs. The docs don't explain this directly, but two details help. The replication point is group-wide: it "shows the last time all in-scope resources were copied together." So an empty catalog only shows as replicated once the whole cycle finishes. Also, the system table doesn't track individual catalogs: it "does not list which individual objects replicated successfully," and a null lag means "at least one asset has never been replicated." To tell whether it's actually stuck or just waiting on the cycle, check the errors column:

 
sql
SELECT event_time, replication_state, replication_lag_ms, errors
FROM system.replication.states
WHERE failover_group_name LIKE '%<your-group>%'
ORDER BY event_time DESC;

The table can take up to 3 hours to populate. If you see a rising lag with no errors, that's worth raising with your account team.
link doc

2. The error about the deleted metastore. The failover group records the metastores it manages (metastore_ids). Swapping the workspace to a new metastore doesn't update the group, so it's still pointing at the one you deleted. The docs don't describe any way to clean that up. There's no force-delete, and nothing on deleting a metastore or workspace that belongs to a failover group. They do say the group sets up a connection and a foreign catalog in each metastore, and that you shouldn't delete those yourself. It's likely that the missing old metastore is what blocks the teardown.
link doc

On deleting the workspace: I wouldn't. The docs never say it cascades to the failover group, so you could end up with an orphaned group, and deleting a workspace can't be undone. Managed DR is gated and enabled by the account team, so they (or a support ticket) are the right channel. Send them the failover group name, the old metastore ID from the error, and the group's current state (probably DELETION_FAILED). This needs a fix on their side.

For future tests, don't reassign or delete a metastore while a failover group references it. The documented teardown is to delete the group first and then turn off Mission Critical on each workspace.

Thomaz A. Rossito Neto
Principal Data Architect & AI Strategy — CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto

Hi Thomas,
I actually tried deleting the workspace and it works! once the workspaces are deleted, the workspace get de-attached from the failover group making failover group empty and then we can delete it successfully.