- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
a week ago
The pattern that keeps this manageable at scale is treating recipients as code and running rotation as a scheduled job. Here's the breakdown:
- Set a token lifetime at the metastore level. Open-sharing bearer tokens expire according to the metastore's recipient token lifetime setting Dont use "no expiry" for external partners.
- Rotate with a grace period. The API supports this directly:
- REST: POST /api/2.1/unity-catalog/recipients/{name}/rotate-token with existing_token_expire_in_seconds
- CLI: databricks recipients rotate-token <name> <existing-token-expire-in-seconds>
- SDK: w.recipients.rotate_token(name=..., existing_token_expire_in_seconds=...)
The old token stays valid for the grace window you set (for example 7 days), so the partner can swap credentials without downtime. The response includes a new activation URL.
- Automate it as a Databricks job. Run a scheduled job that lists recipients (w.recipients.list()) and checks each token's expiration_time. For anything expiring within N days, it rotates the token and sends the new activation link to the partner. Store owner/contact/expiry metadata on the recipient itself with PROPERTIES (e.g. 'partner_contact', 'contract_end') so the job knows whom to notify.
- Deliver the activation link securely. The link is single-use: the credential file can only be downloaded once. That's good for security, but don't paste it into a normal email thread. Put it in a secure transfer channel or a portal the partner authenticates into. If a partner loses the file, rotate again rather than trying to recover it.
- For large programs, consider OIDC federation instead of bearer tokens. Delta Sharing open sharing supports OIDC token federation. Recipients authenticate with their own IdP (Entra ID, Okta, etc.), often machine-to-machine, so you don't have long-lived bearer tokens to rotate at all. It's the cleanest answer to the rotation problem at scale, as long as your partners' tooling supports it.
- Lock down and monitor.
- Attach IP access lists to each recipient where partners have stable egress IPs.
- Monitor usage from system.access.audit: filter service_name = 'unityCatalog' and action_name LIKE 'deltaSharing%' to see which recipients are actually querying. Idle recipients are candidates for offboarding.
- Make offboarding explicit. Use REVOKE SELECT ON SHARE <share> FROM RECIPIENT <recipient> to cut access to specific data, or DROP RECIPIENT to kill all of that recipient's credentials at once. Tie it to the contract end date stored in recipient properties so the same scheduled job can flag or execute it.
To keep this auditable, manage recipients, shares and grants through Terraform or Asset Bundles. Then onboarding and offboarding go through code review instead of UI clicks.
Yes when data is large (usually 50 million + rows), pandas does not perform as good as spark dataframes
Hope this helps 🙂