Hi,
I've tested managed DR as well, and the docs cover less of this than you'd hope, so here's what they do say.
1. Slow replication on empty catalogs. The docs don't explain this directly, but two details help. The replication point is group-wide: it "shows the last time all in-scope resources were copied together." So an empty catalog only shows as replicated once the whole cycle finishes. Also, the system table doesn't track individual catalogs: it "does not list which individual objects replicated successfully," and a null lag means "at least one asset has never been replicated." To tell whether it's actually stuck or just waiting on the cycle, check the errors column:
sql
SELECT event_time, replication_state, replication_lag_ms, errors
FROM system.replication.states
WHERE failover_group_name LIKE '%<your-group>%'
ORDER BY event_time DESC;
The table can take up to 3 hours to populate. If you see a rising lag with no errors, that's worth raising with your account team.
link doc
2. The error about the deleted metastore. The failover group records the metastores it manages (metastore_ids). Swapping the workspace to a new metastore doesn't update the group, so it's still pointing at the one you deleted. The docs don't describe any way to clean that up. There's no force-delete, and nothing on deleting a metastore or workspace that belongs to a failover group. They do say the group sets up a connection and a foreign catalog in each metastore, and that you shouldn't delete those yourself. It's likely that the missing old metastore is what blocks the teardown.
link doc
On deleting the workspace: I wouldn't. The docs never say it cascades to the failover group, so you could end up with an orphaned group, and deleting a workspace can't be undone. Managed DR is gated and enabled by the account team, so they (or a support ticket) are the right channel. Send them the failover group name, the old metastore ID from the error, and the group's current state (probably DELETION_FAILED). This needs a fix on their side.
For future tests, don't reassign or delete a metastore while a failover group references it. The documented teardown is to delete the group first and then turn off Mission Critical on each workspace.
Thomaz A. Rossito Neto
Principal Data Architect & AI Strategy โ CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto