<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Scaling Delta Sharing across Non-Databricks Recipients: Handling Credential Rotation and Lifecyc in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/170023#M56211</link>
    <description>&lt;P&gt;Hi,&lt;/P&gt;&lt;P&gt;1. Credential lifecycle: If your partners have their own identity provider (Entra, Okta, etc.), consider OIDC token federation instead of bearer tokens. The partner's IdP issues short-lived tokens, so there's nothing for you to rotate or leak. It supports both user-based (U2M) and machine-to-machine (M2M) flows.&lt;/P&gt;&lt;P&gt;If you stay on bearer tokens:&lt;/P&gt;&lt;P&gt;Set a metastore-wide token lifetime so tokens expire automatically.&lt;BR /&gt;Automate rotation with a scheduled job that calls the rotate-token API or CLI (databricks recipients rotate-token). Set a short grace period for the old token and send the new activation link to the partner through your own channel.&lt;BR /&gt;Revoke access without manual steps using REVOKE SELECT ON SHARE ... FROM RECIPIENT ... or DROP RECIPIENT, triggered from your offboarding or contract-end process.&lt;BR /&gt;Add IP access lists on each recipient as an extra control.&lt;/P&gt;&lt;P&gt;2. Scale (100+ recipients): Most teams use both:&lt;/P&gt;&lt;P&gt;Terraform (databricks_share, databricks_recipient, databricks_grants) as the source of truth.&lt;BR /&gt;A request workflow in front of it, such as a ServiceNow or Jira form that opens a PR against a recipient config file (YAML/JSON). After approval, CI applies it.&lt;/P&gt;&lt;P&gt;This gives you an audit trail, approvals and easy offboarding. For monitoring, use system.access.audit (Delta Sharing events) to spot unused recipients and clean them up.&lt;/P&gt;&lt;P&gt;3. Spark vs. Python/Pandas: Yes, the difference can be large on big tables:&lt;/P&gt;&lt;P&gt;Spark connector: reads files in parallel, supports partition pruning and predicate pushdown, and handles CDF and streaming. Best for large or incremental reads.&lt;BR /&gt;Pandas connector (load_as_pandas): single-node and loads data into memory. Fine for small or medium tables, but slow or out-of-memory on large ones. Use limit and predicate hints, and newer connector versions that read with delta-kernel.&lt;/P&gt;&lt;P&gt;In practice: partition shared tables on common filter columns, share CDF for incremental pulls, and point pandas users to filtered views or smaller tables.&lt;/P&gt;</description>
    <pubDate>Mon, 28 Sep 2026 09:27:50 GMT</pubDate>
    <dc:creator>Islam_hoti</dc:creator>
    <dc:date>2026-09-28T09:27:50Z</dc:date>
    <item>
      <title>Scaling Delta Sharing across Non-Databricks Recipients: Handling Credential Rotation and Lifecycle</title>
      <link>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/169188#M56050</link>
      <description>&lt;P&gt;Hi everyone,&lt;/P&gt;&lt;P&gt;Hi everyone, I’m architecting a data-sharing program using Delta Sharing for external partners who do not have a Databricks workspace.&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;How are others automating the recipient credential lifecycle? Do you have an automated process for rotating these credentials, or is there a standard pattern for revoking access without manual intervention?&lt;/LI&gt;&lt;LI&gt;For high-scale sharing (100+ recipients), are you maintaining these via Terraform/DABs, or is there a "request-based" workflow you’ve built?&lt;/LI&gt;&lt;LI&gt;Have you encountered significant performance differences between partners using Spark vs. those using native Python/Pandas connectors?&lt;/LI&gt;&lt;/OL&gt;</description>
      <pubDate>Sat, 19 Sep 2026 15:56:19 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/169188#M56050</guid>
      <dc:creator>Khasim_1</dc:creator>
      <dc:date>2026-09-19T15:56:19Z</dc:date>
    </item>
    <item>
      <title>Re: Scaling Delta Sharing across Non-Databricks Recipients: Handling Credential Rotation and Lifecyc</title>
      <link>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/169913#M56190</link>
      <description>&lt;P&gt;The pattern that keeps this manageable at scale is treating recipients as code and running rotation as a scheduled job. Here's the breakdown:&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Set a token lifetime at the metastore level. Open-sharing bearer tokens expire according to the metastore's recipient token lifetime setting Dont use "no expiry" for external partners.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Rotate with a grace period. The API supports this directly:&lt;/P&gt;&lt;P&gt;&amp;nbsp; - REST: POST /api/2.1/unity-catalog/recipients/{name}/rotate-token with existing_token_expire_in_seconds&lt;/P&gt;&lt;P&gt;&amp;nbsp; - CLI: databricks recipients rotate-token &amp;lt;name&amp;gt; &amp;lt;existing-token-expire-in-seconds&amp;gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp; - SDK: w.recipients.rotate_token(name=..., existing_token_expire_in_seconds=...)&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The old token stays valid for the grace window you set (for example 7 days), so the partner can swap credentials without downtime. The response includes a new activation URL.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Automate it as a Databricks job. Run a scheduled job that lists recipients (w.recipients.list()) and checks each token's expiration_time. For anything expiring within N days, it rotates the token and sends the new activation link to the partner. Store owner/contact/expiry metadata on the recipient itself with PROPERTIES (e.g. 'partner_contact', 'contract_end') so the job knows whom to notify.&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Deliver the activation link securely. The link is single-use: the credential file can only be downloaded once. That's good for security, but don't paste it into a normal email thread. Put it in a secure transfer channel or a portal the partner authenticates into. If a partner loses the file, rotate again rather than trying to recover it.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- For large programs, consider OIDC federation instead of bearer tokens. Delta Sharing open sharing supports OIDC token federation. Recipients authenticate with their own IdP (Entra ID, Okta, etc.), often machine-to-machine, so you don't have long-lived bearer tokens to rotate at all. It's the cleanest answer to the rotation problem at scale, as long as your partners' tooling supports it.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Lock down and monitor.&lt;/P&gt;&lt;P&gt;&amp;nbsp; - Attach IP access lists to each recipient where partners have stable egress IPs.&lt;/P&gt;&lt;P&gt;&amp;nbsp; - Monitor usage from system.access.audit: filter service_name = 'unityCatalog' and action_name LIKE 'deltaSharing%' to see which recipients are actually querying. Idle recipients are candidates for offboarding.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Make offboarding explicit. Use REVOKE SELECT ON SHARE &amp;lt;share&amp;gt; FROM RECIPIENT &amp;lt;recipient&amp;gt; to cut access to specific data, or DROP RECIPIENT to kill all of that recipient's credentials at once. Tie it to the contract end date stored in recipient properties so the same scheduled job can flag or execute it.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;To keep this auditable, manage recipients, shares and grants through Terraform or Asset Bundles. Then onboarding and offboarding go through code review instead of UI clicks.&lt;/P&gt;&lt;P&gt;Yes when data is large (usually 50 million + rows), pandas does not perform as good as spark dataframes&lt;/P&gt;&lt;P&gt;Hope this helps &lt;span class="lia-unicode-emoji" title=":slightly_smiling_face:"&gt;🙂&lt;/span&gt;&lt;/P&gt;</description>
      <pubDate>Sat, 26 Sep 2026 16:01:18 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/169913#M56190</guid>
      <dc:creator>SumeshKashyap</dc:creator>
      <dc:date>2026-09-26T16:01:18Z</dc:date>
    </item>
    <item>
      <title>Re: Scaling Delta Sharing across Non-Databricks Recipients: Handling Credential Rotation and Lifecyc</title>
      <link>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/170023#M56211</link>
      <description>&lt;P&gt;Hi,&lt;/P&gt;&lt;P&gt;1. Credential lifecycle: If your partners have their own identity provider (Entra, Okta, etc.), consider OIDC token federation instead of bearer tokens. The partner's IdP issues short-lived tokens, so there's nothing for you to rotate or leak. It supports both user-based (U2M) and machine-to-machine (M2M) flows.&lt;/P&gt;&lt;P&gt;If you stay on bearer tokens:&lt;/P&gt;&lt;P&gt;Set a metastore-wide token lifetime so tokens expire automatically.&lt;BR /&gt;Automate rotation with a scheduled job that calls the rotate-token API or CLI (databricks recipients rotate-token). Set a short grace period for the old token and send the new activation link to the partner through your own channel.&lt;BR /&gt;Revoke access without manual steps using REVOKE SELECT ON SHARE ... FROM RECIPIENT ... or DROP RECIPIENT, triggered from your offboarding or contract-end process.&lt;BR /&gt;Add IP access lists on each recipient as an extra control.&lt;/P&gt;&lt;P&gt;2. Scale (100+ recipients): Most teams use both:&lt;/P&gt;&lt;P&gt;Terraform (databricks_share, databricks_recipient, databricks_grants) as the source of truth.&lt;BR /&gt;A request workflow in front of it, such as a ServiceNow or Jira form that opens a PR against a recipient config file (YAML/JSON). After approval, CI applies it.&lt;/P&gt;&lt;P&gt;This gives you an audit trail, approvals and easy offboarding. For monitoring, use system.access.audit (Delta Sharing events) to spot unused recipients and clean them up.&lt;/P&gt;&lt;P&gt;3. Spark vs. Python/Pandas: Yes, the difference can be large on big tables:&lt;/P&gt;&lt;P&gt;Spark connector: reads files in parallel, supports partition pruning and predicate pushdown, and handles CDF and streaming. Best for large or incremental reads.&lt;BR /&gt;Pandas connector (load_as_pandas): single-node and loads data into memory. Fine for small or medium tables, but slow or out-of-memory on large ones. Use limit and predicate hints, and newer connector versions that read with delta-kernel.&lt;/P&gt;&lt;P&gt;In practice: partition shared tables on common filter columns, share CDF for incremental pulls, and point pandas users to filtered views or smaller tables.&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 09:27:50 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/scaling-delta-sharing-across-non-databricks-recipients-handling/m-p/170023#M56211</guid>
      <dc:creator>Islam_hoti</dc:creator>
      <dc:date>2026-09-28T09:27:50Z</dc:date>
    </item>
  </channel>
</rss>

