<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Azure to AWS in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170448#M56299</link>
    <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262067"&gt;@priya9896&lt;/a&gt;, I took a look at both internal and external documentation and here is what I found.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Straight answer:&lt;/STRONG&gt; don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itself are built around tables in S3. Even Glue catalog federation to Unity Catalog's Iceberg REST endpoint works by having Lake Formation vend scoped credentials to data stored in S3, and AWS's walkthrough for it lists a Databricks workspace on AWS as a prerequisite. So: can a Glue Spark job read Azure Databricks data without copying it? Yes. Can you get Glue Catalog tables that Athena queries without a copy? No. Your two goals pull against each other, and one has to give.&lt;/P&gt;
&lt;P&gt;With "no copy" as the hard constraint, the supported pattern is Delta Sharing, which Databricks renamed OpenSharing in June 2026. Same protocol; the package is still &lt;CODE&gt;delta-sharing-spark&lt;/CODE&gt; and the format is still &lt;CODE&gt;deltaSharing&lt;/CODE&gt;.&lt;/P&gt;
&lt;P&gt;If your AWS consumers have a Unity Catalog-enabled Databricks workspace, use Databricks-to-Databricks sharing. The shared tables show up read-only in their catalog and there's no credential file to manage.&lt;/P&gt;
&lt;P&gt;If Glue is the actual consumer, use Databricks-to-Open sharing. The source team creates a share with only the tables you need (partition filters or shared views if they want to narrow rows or columns), creates a recipient (bearer token or OIDC), and sends you a credential file. A Glue Spark job then reads it like this, with the credential file locked down since it is a bearer token:&lt;/P&gt;
&lt;P&gt;&lt;CODE&gt;df = spark.read.format("deltaSharing").load("s3://&amp;lt;secure-bucket&amp;gt;/config.share#&amp;lt;share&amp;gt;.&amp;lt;schema&amp;gt;.&amp;lt;table&amp;gt;")&lt;/CODE&gt;&lt;/P&gt;
&lt;P&gt;The same credential file carries an &lt;CODE&gt;icebergEndpoint&lt;/CODE&gt;, so Iceberg clients (Spark with an Iceberg REST catalog, PyIceberg, Trino, Snowflake) can read the share too.&lt;/P&gt;
&lt;P&gt;Either way, Databricks returns the table's storage location with temporary cloud credentials, and your compute reads directly from Azure storage. Nothing persists in AWS. That's no replication, but it isn't zero network movement: rows cross the cloud boundary on every query. Prototype it from a Glue job, measure latency and throughput, and settle egress with the source team up front, because your cloud vendor may charge egress fees for sharing across clouds, and those land on whoever owns the storage, which is them. Databricks documents Cloudflare R2 as an egress-free option (a replica, so their call), and with SecureConnect, Databricks bills the data transfer rather than the cloud vendor; it's in Public Preview, so ask about it. Also pin down what "no ingestion into AWS" means. If it means no persisted copy, OpenSharing fits. If it means no bytes leave Azure, nothing here works and the consumers need to compute in Azure.&lt;/P&gt;
&lt;P&gt;On the other patterns:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Unity Catalog's Iceberg REST catalog or Unity REST API with a service principal. Works from Spark, but the source team has to enable external data access on their metastore and manage a principal for you. OpenSharing is the purpose-built version for a separate consuming team: scoped to a share, revocable, audited.&lt;/LI&gt;
&lt;LI&gt;JDBC to a Databricks SQL warehouse. Fine for targeted extracts, but every read runs on the source team's warehouse, rows move over JDBC rather than parallel Parquet reads, and the result still isn't a native Glue Catalog table Athena can query.&lt;/LI&gt;
&lt;LI&gt;Direct storage access. AWS does publish an Azure Data Lake Storage connector for Glue that can read Delta tables straight from ADLS Gen2, but it runs on storage credentials, not Unity Catalog permissions, so it bypasses the source team's governance entirely. Most source teams won't agree, and I wouldn't ask.&lt;/LI&gt;
&lt;LI&gt;Lakehouse Federation runs the other direction (Databricks querying external systems), so it doesn't help here.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If your consumers truly need &lt;CODE&gt;glueContext.create_data_frame.from_catalog(...)&lt;/CODE&gt; or Athena, that's a different problem, and the only route is a governed, scheduled replica into S3 that Glue then catalogs. That's replication, which the source team has ruled out, so I'd have that conversation with them rather than engineer around it. First, though, confirm whether the consumers need Glue Catalog metadata or just governed read access. In my experience it's usually the latter, and OpenSharing covers it.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;OpenSharing announcement (Delta Sharing rename): &lt;A href="https://www.databricks.com/blog/announcing-new-opensharing-and-marketplace-capabilities-ai-era" target="_blank"&gt;https://www.databricks.com/blog/announcing-new-opensharing-and-marketplace-capabilities-ai-era&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;OpenSharing overview (Azure Databricks): &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks-to-Open sharing, provider side: &lt;A href="https://learn.microsoft.com/azure/databricks/opensharing/share-data-open" target="_blank"&gt;https://learn.microsoft.com/azure/databricks/opensharing/share-data-open&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Read shared data with Spark and Iceberg clients, recipient side: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/read-data-open" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/read-data-open&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Monitor and manage OpenSharing egress costs: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Unity Catalog Iceberg REST catalog for external engines: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/external-access/iceberg" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/external-access/iceberg&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Delta Sharing open-source Spark connector: &lt;A href="https://docs.delta.io/delta-sharing/" target="_blank"&gt;https://docs.delta.io/delta-sharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue catalog federation to Unity Catalog: &lt;A href="https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html" target="_blank"&gt;https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS blog, Unity Catalog federation walkthrough: &lt;A href="https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/" target="_blank"&gt;https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue Delta Lake framework: &lt;A href="https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format-delta-lake.html" target="_blank"&gt;https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format-delta-lake.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS blog, reading Delta tables from ADLS Gen2 with the Glue connector: &lt;A href="https://aws.amazon.com/blogs/big-data/migrate-delta-tables-from-azure-data-lake-storage-to-amazon-s3-using-aws-glue/" target="_blank"&gt;https://aws.amazon.com/blogs/big-data/migrate-delta-tables-from-azure-data-lake-storage-to-amazon-s3-using-aws-glue/&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
    <pubDate>Fri, 02 Oct 2026 17:22:04 GMT</pubDate>
    <dc:creator>Louis_Frolio</dc:creator>
    <dc:date>2026-10-02T17:22:04Z</dc:date>
    <item>
      <title>Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170186#M56247</link>
      <description>&lt;P class=""&gt;Hi everyone,&lt;/P&gt;&lt;P class=""&gt;We're evaluating an architecture pattern and would appreciate any guidance or recommendations.&lt;/P&gt;&lt;P class=""&gt;&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;Has anyone implemented a similar cross-cloud pattern? Specifically:&lt;/P&gt;&lt;UL class=""&gt;&lt;LI&gt;Can AWS Glue be used to access Azure Databricks data without copying it?&lt;/LI&gt;&lt;LI&gt;Are there recommended approaches such as federation, Delta Sharing, custom connectors, JDBC/SQL endpoints, or other patterns?&lt;/LI&gt;&lt;/UL&gt;&lt;P class=""&gt;Thanks in advance!&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 19:17:54 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170186#M56247</guid>
      <dc:creator>priya9896</dc:creator>
      <dc:date>2026-09-29T19:17:54Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170187#M56248</link>
      <description>&lt;P&gt;Hi everyone,&lt;/P&gt;&lt;P&gt;We're evaluating an architecture pattern and would appreciate any guidance or recommendations.&lt;/P&gt;&lt;P&gt;Current state:&lt;/P&gt;&lt;P&gt;Source data resides in Azure Databricks.&lt;BR /&gt;Consumers are in AWS .&lt;BR /&gt;The source team allows read-only access but does not allow data replication or ingestion into AWS.&lt;BR /&gt;Goal: We would like AWS consumers to access the data through Glue Catalog tables while keeping the data in Azure Databricks.&lt;/P&gt;&lt;P&gt;Has anyone implemented a similar cross-cloud pattern? Specifically:&lt;/P&gt;&lt;P&gt;Can AWS Glue be used to access Azure Databricks data without copying it?&lt;BR /&gt;Are there recommended approaches such as federation, Delta Sharing, custom connectors, JDBC/SQL endpoints, or other patterns?&lt;BR /&gt;&lt;BR /&gt;Appreciate all your time and inputs regarding . Thanks in advance!&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 19:21:38 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170187#M56248</guid>
      <dc:creator>priya9896</dc:creator>
      <dc:date>2026-09-29T19:21:38Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170448#M56299</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262067"&gt;@priya9896&lt;/a&gt;, I took a look at both internal and external documentation and here is what I found.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Straight answer:&lt;/STRONG&gt; don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itself are built around tables in S3. Even Glue catalog federation to Unity Catalog's Iceberg REST endpoint works by having Lake Formation vend scoped credentials to data stored in S3, and AWS's walkthrough for it lists a Databricks workspace on AWS as a prerequisite. So: can a Glue Spark job read Azure Databricks data without copying it? Yes. Can you get Glue Catalog tables that Athena queries without a copy? No. Your two goals pull against each other, and one has to give.&lt;/P&gt;
&lt;P&gt;With "no copy" as the hard constraint, the supported pattern is Delta Sharing, which Databricks renamed OpenSharing in June 2026. Same protocol; the package is still &lt;CODE&gt;delta-sharing-spark&lt;/CODE&gt; and the format is still &lt;CODE&gt;deltaSharing&lt;/CODE&gt;.&lt;/P&gt;
&lt;P&gt;If your AWS consumers have a Unity Catalog-enabled Databricks workspace, use Databricks-to-Databricks sharing. The shared tables show up read-only in their catalog and there's no credential file to manage.&lt;/P&gt;
&lt;P&gt;If Glue is the actual consumer, use Databricks-to-Open sharing. The source team creates a share with only the tables you need (partition filters or shared views if they want to narrow rows or columns), creates a recipient (bearer token or OIDC), and sends you a credential file. A Glue Spark job then reads it like this, with the credential file locked down since it is a bearer token:&lt;/P&gt;
&lt;P&gt;&lt;CODE&gt;df = spark.read.format("deltaSharing").load("s3://&amp;lt;secure-bucket&amp;gt;/config.share#&amp;lt;share&amp;gt;.&amp;lt;schema&amp;gt;.&amp;lt;table&amp;gt;")&lt;/CODE&gt;&lt;/P&gt;
&lt;P&gt;The same credential file carries an &lt;CODE&gt;icebergEndpoint&lt;/CODE&gt;, so Iceberg clients (Spark with an Iceberg REST catalog, PyIceberg, Trino, Snowflake) can read the share too.&lt;/P&gt;
&lt;P&gt;Either way, Databricks returns the table's storage location with temporary cloud credentials, and your compute reads directly from Azure storage. Nothing persists in AWS. That's no replication, but it isn't zero network movement: rows cross the cloud boundary on every query. Prototype it from a Glue job, measure latency and throughput, and settle egress with the source team up front, because your cloud vendor may charge egress fees for sharing across clouds, and those land on whoever owns the storage, which is them. Databricks documents Cloudflare R2 as an egress-free option (a replica, so their call), and with SecureConnect, Databricks bills the data transfer rather than the cloud vendor; it's in Public Preview, so ask about it. Also pin down what "no ingestion into AWS" means. If it means no persisted copy, OpenSharing fits. If it means no bytes leave Azure, nothing here works and the consumers need to compute in Azure.&lt;/P&gt;
&lt;P&gt;On the other patterns:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Unity Catalog's Iceberg REST catalog or Unity REST API with a service principal. Works from Spark, but the source team has to enable external data access on their metastore and manage a principal for you. OpenSharing is the purpose-built version for a separate consuming team: scoped to a share, revocable, audited.&lt;/LI&gt;
&lt;LI&gt;JDBC to a Databricks SQL warehouse. Fine for targeted extracts, but every read runs on the source team's warehouse, rows move over JDBC rather than parallel Parquet reads, and the result still isn't a native Glue Catalog table Athena can query.&lt;/LI&gt;
&lt;LI&gt;Direct storage access. AWS does publish an Azure Data Lake Storage connector for Glue that can read Delta tables straight from ADLS Gen2, but it runs on storage credentials, not Unity Catalog permissions, so it bypasses the source team's governance entirely. Most source teams won't agree, and I wouldn't ask.&lt;/LI&gt;
&lt;LI&gt;Lakehouse Federation runs the other direction (Databricks querying external systems), so it doesn't help here.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If your consumers truly need &lt;CODE&gt;glueContext.create_data_frame.from_catalog(...)&lt;/CODE&gt; or Athena, that's a different problem, and the only route is a governed, scheduled replica into S3 that Glue then catalogs. That's replication, which the source team has ruled out, so I'd have that conversation with them rather than engineer around it. First, though, confirm whether the consumers need Glue Catalog metadata or just governed read access. In my experience it's usually the latter, and OpenSharing covers it.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;OpenSharing announcement (Delta Sharing rename): &lt;A href="https://www.databricks.com/blog/announcing-new-opensharing-and-marketplace-capabilities-ai-era" target="_blank"&gt;https://www.databricks.com/blog/announcing-new-opensharing-and-marketplace-capabilities-ai-era&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;OpenSharing overview (Azure Databricks): &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks-to-Open sharing, provider side: &lt;A href="https://learn.microsoft.com/azure/databricks/opensharing/share-data-open" target="_blank"&gt;https://learn.microsoft.com/azure/databricks/opensharing/share-data-open&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Read shared data with Spark and Iceberg clients, recipient side: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/read-data-open" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/read-data-open&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Monitor and manage OpenSharing egress costs: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Unity Catalog Iceberg REST catalog for external engines: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/external-access/iceberg" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/external-access/iceberg&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Delta Sharing open-source Spark connector: &lt;A href="https://docs.delta.io/delta-sharing/" target="_blank"&gt;https://docs.delta.io/delta-sharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue catalog federation to Unity Catalog: &lt;A href="https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html" target="_blank"&gt;https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS blog, Unity Catalog federation walkthrough: &lt;A href="https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/" target="_blank"&gt;https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue Delta Lake framework: &lt;A href="https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format-delta-lake.html" target="_blank"&gt;https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format-delta-lake.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS blog, reading Delta tables from ADLS Gen2 with the Glue connector: &lt;A href="https://aws.amazon.com/blogs/big-data/migrate-delta-tables-from-azure-data-lake-storage-to-amazon-s3-using-aws-glue/" target="_blank"&gt;https://aws.amazon.com/blogs/big-data/migrate-delta-tables-from-azure-data-lake-storage-to-amazon-s3-using-aws-glue/&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Fri, 02 Oct 2026 17:22:04 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170448#M56299</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-02T17:22:04Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170466#M56306</link>
      <description>&lt;P&gt;Great explanation! I'd add one important consideration from a cross-cloud architecture and data governance perspective.&lt;/P&gt;&lt;P&gt;Before implementing OpenSharing with AWS Glue, I would clarify exactly what the source team means by "no replication or ingestion into AWS."&lt;/P&gt;&lt;P&gt;There is an important difference between avoiding a permanent replica in S3 and prohibiting any temporary data persistence or processing within AWS.&lt;/P&gt;&lt;P&gt;Even when using OpenSharing, the data still crosses cloud boundaries. Depending on the Glue job configuration, intermediate Spark data could also be written to temporary storage.&lt;/P&gt;&lt;P&gt;I'd suggest validating three things during a small proof of concept:&lt;/P&gt;&lt;P&gt;Governance: Confirm whether temporary processing and potential data spilling in AWS are permitted.&lt;/P&gt;&lt;P&gt;Performance: Measure cross-cloud transfer costs, query latency and the volume of data transferred.&lt;/P&gt;&lt;P&gt;Operations: Validate credentials, network connectivity and how shared-table schema changes affect downstream consumers.&lt;/P&gt;&lt;P&gt;If the requirement is strictly read-only access without a permanent replica, OpenSharing is worth evaluating.&lt;/P&gt;&lt;P&gt;However, if no data is permitted to leave Azure, I would consider keeping the processing in Azure and exposing only approved results.&lt;/P&gt;&lt;P&gt;One question: Is the Glue Catalog requirement driven by Athena integration, or could the AWS consumers work directly with shared datasets through Glue Spark?&lt;/P&gt;</description>
      <pubDate>Fri, 02 Oct 2026 22:43:41 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170466#M56306</guid>
      <dc:creator>lmcorreahdb</dc:creator>
      <dc:date>2026-10-02T22:43:41Z</dc:date>
    </item>
  </channel>
</rss>

