<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Azure to AWS in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172195#M56593</link>
    <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262067"&gt;@priya9896&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Given the no-replication requirement, I would evaluate Delta Sharing/OpenSharing first rather than JDBC-based ingestion. It is designed for secure, read-only cross-platform data sharing without maintaining a replicated copy of the source data. &lt;A href="https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-sharing?utm_source=chatgpt.com" target="_blank"&gt;Databricks Documentation&lt;/A&gt;&lt;/P&gt;&lt;P&gt;The pattern would be:&lt;/P&gt;&lt;P&gt;Azure Databricks / Unity Catalog → Delta Sharing/OpenSharing → AWS consumer&lt;/P&gt;&lt;P&gt;The important distinction is that Glue Catalog is a metadata/catalog layer, so I wouldn’t treat “creating a Glue table” itself as the mechanism for accessing the Azure Databricks data. The AWS-side engine still needs a supported way to consume the shared data.&lt;/P&gt;&lt;P&gt;I would therefore validate which AWS compute engine needs to query those Glue tables first, and then determine whether it can consume the OpenSharing share directly or requires an integration layer.&lt;/P&gt;&lt;P&gt;Also consider cross-cloud network egress cost and network controls, even though the architecture avoids maintaining another copy of the dataset.&lt;/P&gt;</description>
    <pubDate>Wed, 07 Oct 2026 18:18:33 GMT</pubDate>
    <dc:creator>vs4</dc:creator>
    <dc:date>2026-10-07T18:18:33Z</dc:date>
    <item>
      <title>Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170186#M56247</link>
      <description>&lt;P class=""&gt;Hi everyone,&lt;/P&gt;&lt;P class=""&gt;We're evaluating an architecture pattern and would appreciate any guidance or recommendations.&lt;/P&gt;&lt;P class=""&gt;&amp;nbsp;&lt;/P&gt;&lt;P class=""&gt;Has anyone implemented a similar cross-cloud pattern? Specifically:&lt;/P&gt;&lt;UL class=""&gt;&lt;LI&gt;Can AWS Glue be used to access Azure Databricks data without copying it?&lt;/LI&gt;&lt;LI&gt;Are there recommended approaches such as federation, Delta Sharing, custom connectors, JDBC/SQL endpoints, or other patterns?&lt;/LI&gt;&lt;/UL&gt;&lt;P class=""&gt;Thanks in advance!&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 19:17:54 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170186#M56247</guid>
      <dc:creator>priya9896</dc:creator>
      <dc:date>2026-09-29T19:17:54Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170187#M56248</link>
      <description>&lt;P&gt;Hi everyone,&lt;/P&gt;&lt;P&gt;We're evaluating an architecture pattern and would appreciate any guidance or recommendations.&lt;/P&gt;&lt;P&gt;Current state:&lt;/P&gt;&lt;P&gt;Source data resides in Azure Databricks.&lt;BR /&gt;Consumers are in AWS .&lt;BR /&gt;The source team allows read-only access but does not allow data replication or ingestion into AWS.&lt;BR /&gt;Goal: We would like AWS consumers to access the data through Glue Catalog tables while keeping the data in Azure Databricks.&lt;/P&gt;&lt;P&gt;Has anyone implemented a similar cross-cloud pattern? Specifically:&lt;/P&gt;&lt;P&gt;Can AWS Glue be used to access Azure Databricks data without copying it?&lt;BR /&gt;Are there recommended approaches such as federation, Delta Sharing, custom connectors, JDBC/SQL endpoints, or other patterns?&lt;BR /&gt;&lt;BR /&gt;Appreciate all your time and inputs regarding . Thanks in advance!&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 19:21:38 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170187#M56248</guid>
      <dc:creator>priya9896</dc:creator>
      <dc:date>2026-09-29T19:21:38Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170448#M56299</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262067"&gt;@priya9896&lt;/a&gt;, I took a look at both internal and external documentation and here is what I found.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Straight answer:&lt;/STRONG&gt; don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itself are built around tables in S3. Even Glue catalog federation to Unity Catalog's Iceberg REST endpoint works by having Lake Formation vend scoped credentials to data stored in S3, and AWS's walkthrough for it lists a Databricks workspace on AWS as a prerequisite. So: can a Glue Spark job read Azure Databricks data without copying it? Yes. Can you get Glue Catalog tables that Athena queries without a copy? No. Your two goals pull against each other, and one has to give.&lt;/P&gt;
&lt;P&gt;With "no copy" as the hard constraint, the supported pattern is Delta Sharing, which Databricks renamed OpenSharing in June 2026. Same protocol; the package is still &lt;CODE&gt;delta-sharing-spark&lt;/CODE&gt; and the format is still &lt;CODE&gt;deltaSharing&lt;/CODE&gt;.&lt;/P&gt;
&lt;P&gt;If your AWS consumers have a Unity Catalog-enabled Databricks workspace, use Databricks-to-Databricks sharing. The shared tables show up read-only in their catalog and there's no credential file to manage.&lt;/P&gt;
&lt;P&gt;If Glue is the actual consumer, use Databricks-to-Open sharing. The source team creates a share with only the tables you need (partition filters or shared views if they want to narrow rows or columns), creates a recipient (bearer token or OIDC), and sends you a credential file. A Glue Spark job then reads it like this, with the credential file locked down since it is a bearer token:&lt;/P&gt;
&lt;P&gt;&lt;CODE&gt;df = spark.read.format("deltaSharing").load("s3://&amp;lt;secure-bucket&amp;gt;/config.share#&amp;lt;share&amp;gt;.&amp;lt;schema&amp;gt;.&amp;lt;table&amp;gt;")&lt;/CODE&gt;&lt;/P&gt;
&lt;P&gt;The same credential file carries an &lt;CODE&gt;icebergEndpoint&lt;/CODE&gt;, so Iceberg clients (Spark with an Iceberg REST catalog, PyIceberg, Trino, Snowflake) can read the share too.&lt;/P&gt;
&lt;P&gt;Either way, Databricks returns the table's storage location with temporary cloud credentials, and your compute reads directly from Azure storage. Nothing persists in AWS. That's no replication, but it isn't zero network movement: rows cross the cloud boundary on every query. Prototype it from a Glue job, measure latency and throughput, and settle egress with the source team up front, because your cloud vendor may charge egress fees for sharing across clouds, and those land on whoever owns the storage, which is them. Databricks documents Cloudflare R2 as an egress-free option (a replica, so their call), and with SecureConnect, Databricks bills the data transfer rather than the cloud vendor; it's in Public Preview, so ask about it. Also pin down what "no ingestion into AWS" means. If it means no persisted copy, OpenSharing fits. If it means no bytes leave Azure, nothing here works and the consumers need to compute in Azure.&lt;/P&gt;
&lt;P&gt;On the other patterns:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Unity Catalog's Iceberg REST catalog or Unity REST API with a service principal. Works from Spark, but the source team has to enable external data access on their metastore and manage a principal for you. OpenSharing is the purpose-built version for a separate consuming team: scoped to a share, revocable, audited.&lt;/LI&gt;
&lt;LI&gt;JDBC to a Databricks SQL warehouse. Fine for targeted extracts, but every read runs on the source team's warehouse, rows move over JDBC rather than parallel Parquet reads, and the result still isn't a native Glue Catalog table Athena can query.&lt;/LI&gt;
&lt;LI&gt;Direct storage access. AWS does publish an Azure Data Lake Storage connector for Glue that can read Delta tables straight from ADLS Gen2, but it runs on storage credentials, not Unity Catalog permissions, so it bypasses the source team's governance entirely. Most source teams won't agree, and I wouldn't ask.&lt;/LI&gt;
&lt;LI&gt;Lakehouse Federation runs the other direction (Databricks querying external systems), so it doesn't help here.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If your consumers truly need &lt;CODE&gt;glueContext.create_data_frame.from_catalog(...)&lt;/CODE&gt; or Athena, that's a different problem, and the only route is a governed, scheduled replica into S3 that Glue then catalogs. That's replication, which the source team has ruled out, so I'd have that conversation with them rather than engineer around it. First, though, confirm whether the consumers need Glue Catalog metadata or just governed read access. In my experience it's usually the latter, and OpenSharing covers it.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;OpenSharing announcement (Delta Sharing rename): &lt;A href="https://www.databricks.com/blog/announcing-new-opensharing-and-marketplace-capabilities-ai-era" target="_blank"&gt;https://www.databricks.com/blog/announcing-new-opensharing-and-marketplace-capabilities-ai-era&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;OpenSharing overview (Azure Databricks): &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks-to-Open sharing, provider side: &lt;A href="https://learn.microsoft.com/azure/databricks/opensharing/share-data-open" target="_blank"&gt;https://learn.microsoft.com/azure/databricks/opensharing/share-data-open&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Read shared data with Spark and Iceberg clients, recipient side: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/read-data-open" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/read-data-open&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Monitor and manage OpenSharing egress costs: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Unity Catalog Iceberg REST catalog for external engines: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/external-access/iceberg" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/external-access/iceberg&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Delta Sharing open-source Spark connector: &lt;A href="https://docs.delta.io/delta-sharing/" target="_blank"&gt;https://docs.delta.io/delta-sharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue catalog federation to Unity Catalog: &lt;A href="https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html" target="_blank"&gt;https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS blog, Unity Catalog federation walkthrough: &lt;A href="https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/" target="_blank"&gt;https://aws.amazon.com/blogs/big-data/access-databricks-unity-catalog-data-using-catalog-federation-in-the-aws-glue-data-catalog/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue Delta Lake framework: &lt;A href="https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format-delta-lake.html" target="_blank"&gt;https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-format-delta-lake.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS blog, reading Delta tables from ADLS Gen2 with the Glue connector: &lt;A href="https://aws.amazon.com/blogs/big-data/migrate-delta-tables-from-azure-data-lake-storage-to-amazon-s3-using-aws-glue/" target="_blank"&gt;https://aws.amazon.com/blogs/big-data/migrate-delta-tables-from-azure-data-lake-storage-to-amazon-s3-using-aws-glue/&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Fri, 02 Oct 2026 17:22:04 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170448#M56299</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-02T17:22:04Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170466#M56306</link>
      <description>&lt;P&gt;Great explanation! I'd add one important consideration from a cross-cloud architecture and data governance perspective.&lt;/P&gt;&lt;P&gt;Before implementing OpenSharing with AWS Glue, I would clarify exactly what the source team means by "no replication or ingestion into AWS."&lt;/P&gt;&lt;P&gt;There is an important difference between avoiding a permanent replica in S3 and prohibiting any temporary data persistence or processing within AWS.&lt;/P&gt;&lt;P&gt;Even when using OpenSharing, the data still crosses cloud boundaries. Depending on the Glue job configuration, intermediate Spark data could also be written to temporary storage.&lt;/P&gt;&lt;P&gt;I'd suggest validating three things during a small proof of concept:&lt;/P&gt;&lt;P&gt;Governance: Confirm whether temporary processing and potential data spilling in AWS are permitted.&lt;/P&gt;&lt;P&gt;Performance: Measure cross-cloud transfer costs, query latency and the volume of data transferred.&lt;/P&gt;&lt;P&gt;Operations: Validate credentials, network connectivity and how shared-table schema changes affect downstream consumers.&lt;/P&gt;&lt;P&gt;If the requirement is strictly read-only access without a permanent replica, OpenSharing is worth evaluating.&lt;/P&gt;&lt;P&gt;However, if no data is permitted to leave Azure, I would consider keeping the processing in Azure and exposing only approved results.&lt;/P&gt;&lt;P&gt;One question: Is the Glue Catalog requirement driven by Athena integration, or could the AWS consumers work directly with shared datasets through Glue Spark?&lt;/P&gt;</description>
      <pubDate>Fri, 02 Oct 2026 22:43:41 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/170466#M56306</guid>
      <dc:creator>lmcorreahdb</dc:creator>
      <dc:date>2026-10-02T22:43:41Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172164#M56587</link>
      <description>&lt;P&gt;Greetings &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262067"&gt;@priya9896&lt;/a&gt;, I looked into things a bit more and here is what I found. And thanks to &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/248665"&gt;@lmcorreahdb&lt;/a&gt; for the governance point; the definition question is the whole ballgame, so most of this puts specifics behind it.&lt;/P&gt;
&lt;P&gt;First, a sharper version of the Glue federation point from my earlier reply. AWS now documents a Glue/Lake Formation federated catalog to Unity Catalog (connection type &lt;CODE&gt;DATABRICKSICEBERGRESTCATALOG&lt;/CODE&gt;), and in principle you could point it at the Azure workspace and see the catalog's databases and tables in Glue. Queries would fail, though. AWS states that only Iceberg tables in Amazon S3 are queryable; any other storage lists fine and then fails at &lt;CODE&gt;SELECT&lt;/CODE&gt; with "Object storage location is not supported." So the Glue Catalog door only opens if the tables live in S3 as Iceberg (or Delta with UniForm), which for an Azure source means a copy. Same conclusion as before, now with the exact error you'd hit.&lt;/P&gt;
&lt;P&gt;Where bytes can land in AWS during a Glue job, even with OpenSharing and no explicit write:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Shuffle and spill. By default they go to the Glue workers' local disks and vanish with the job. If someone enables &lt;CODE&gt;--write-shuffle-files-to-s3&lt;/CODE&gt;, shuffle files go to an S3 bucket, and AWS notes the shuffle manager doesn't clean them up afterward. Leave that flag off here.&lt;/LI&gt;
&lt;LI&gt;&lt;CODE&gt;--TempDir&lt;/CODE&gt;. Glue's S3 scratch location; some connectors and operations stage data there.&lt;/LI&gt;
&lt;LI&gt;Logs. Continuous CloudWatch logging and Spark UI event logs (&lt;CODE&gt;--enable-spark-ui&lt;/CODE&gt;, &lt;CODE&gt;--spark-event-logs-path&lt;/CODE&gt;) carry driver and executor output and query plans. No rows unless somebody prints a DataFrame, so ban &lt;CODE&gt;show()&lt;/CODE&gt; and &lt;CODE&gt;collect()&lt;/CODE&gt; in production jobs.&lt;/LI&gt;
&lt;LI&gt;Developer habits. &lt;CODE&gt;cache()&lt;/CODE&gt; to disk, checkpoints, a "temporary" debug write. Policy catches these, tooling doesn't.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Write those down and ask the source team which are acceptable. That turns "no ingestion" from a slogan into a spec.&lt;/P&gt;
&lt;P&gt;On the provider side, Databricks-to-Open sharing gives the source team more control than they probably expect:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Scope. The recipient sees only the tables and views in the share, with partition filters to narrow rows. Nothing else in the metastore is visible.&lt;/LI&gt;
&lt;LI&gt;Token lifetime. Set at the metastore or per recipient, rotated on demand; Databricks recommends rotating or dropping a recipient promptly once its token expires.&lt;/LI&gt;
&lt;LI&gt;IP access lists. An allow list on the recipient covering the sharing API, activation link and credential download. Without SecureConnect the short-lived storage URLs can still be used from any IP; with it the list applies to storage too. Run the Glue job through a VPC connection with a NAT gateway if you want stable egress IPs for that list.&lt;/LI&gt;
&lt;LI&gt;Audit. Recipient queries against the share are logged in the provider's &lt;CODE&gt;system.access.audit&lt;/CODE&gt;, so the source team sees who read what and when.&lt;/LI&gt;
&lt;LI&gt;Schema changes. They show up at the recipient's next read. Select explicit columns rather than &lt;CODE&gt;*&lt;/CODE&gt; and agree on a heads-up process with the provider.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;On the Athena question &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/248665"&gt;@lmcorreahdb&lt;/a&gt; raised: if Glue Catalog really means Athena, close that door early. AWS's prebuilt Athena federated connectors don't include Databricks, the generic JDBC connector is deprecated, and the one Azure-oriented connector they publish (ADLS Gen2) routes through Azure Synapse rather than Databricks, so it sidesteps Unity Catalog, and like every Athena connector it spills results to an S3 bucket when a response outgrows Lambda's limits. Athena means a replica, plain and simple.&lt;/P&gt;
&lt;P&gt;So the decision tree is short. Glue Spark (or EMR, SageMaker, anything running the sharing connector) with no permanent replica: OpenSharing, Databricks-to-Open. A Unity Catalog-enabled Databricks workspace on your side: Databricks-to-Databricks, simpler still. Glue Catalog or Athena mandatory: a governed S3 replica, and have that conversation now rather than after the POC. No bytes may leave Azure: compute in Azure and expose approved results only.&lt;/P&gt;
&lt;P&gt;For the POC, one small table from a Glue job, shuffle-to-S3 off, nothing written, then have the source team find the reads in &lt;CODE&gt;system.access.audit&lt;/CODE&gt;. Seeing their own audit trail does more for the governance conversation than any diagram. Please report back with the latency and egress numbers; plenty of folks will hit this same wall.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;AWS Lake Formation, federate to Databricks Unity Catalog (S3-only query note): &lt;A href="https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html" target="_blank"&gt;https://docs.aws.amazon.com/lake-formation/latest/dg/catalog-federation-databricks.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue job parameters (&lt;CODE&gt;--TempDir&lt;/CODE&gt;, Spark UI logs, continuous logging): &lt;A href="https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html" target="_blank"&gt;https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;AWS Glue Spark shuffle plugin with Amazon S3: &lt;A href="https://docs.aws.amazon.com/glue/latest/dg/monitor-spark-shuffle-manager.html" target="_blank"&gt;https://docs.aws.amazon.com/glue/latest/dg/monitor-spark-shuffle-manager.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Create a bearer-token recipient, token lifetime and rotation (Azure Databricks): &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/create-recipient-token" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/create-recipient-token&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Restrict recipient access with IP access lists: &lt;A href="https://learn.microsoft.com/azure/databricks/opensharing/access-list" target="_blank"&gt;https://learn.microsoft.com/azure/databricks/opensharing/access-list&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Audit and monitor data sharing: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/audit-logs" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/audit-logs&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Athena prebuilt data source connectors: &lt;A href="https://docs.aws.amazon.com/athena/latest/ug/connectors-available.html" target="_blank"&gt;https://docs.aws.amazon.com/athena/latest/ug/connectors-available.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Athena ADLS Gen2 connector: &lt;A href="https://docs.aws.amazon.com/athena/latest/ug/connectors-adls-gen2.html" target="_blank"&gt;https://docs.aws.amazon.com/athena/latest/ug/connectors-adls-gen2.html&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Wed, 07 Oct 2026 14:10:33 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172164#M56587</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-07T14:10:33Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172195#M56593</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262067"&gt;@priya9896&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Given the no-replication requirement, I would evaluate Delta Sharing/OpenSharing first rather than JDBC-based ingestion. It is designed for secure, read-only cross-platform data sharing without maintaining a replicated copy of the source data. &lt;A href="https://docs.databricks.com/aws/en/sql/language-manual/sql-ref-sharing?utm_source=chatgpt.com" target="_blank"&gt;Databricks Documentation&lt;/A&gt;&lt;/P&gt;&lt;P&gt;The pattern would be:&lt;/P&gt;&lt;P&gt;Azure Databricks / Unity Catalog → Delta Sharing/OpenSharing → AWS consumer&lt;/P&gt;&lt;P&gt;The important distinction is that Glue Catalog is a metadata/catalog layer, so I wouldn’t treat “creating a Glue table” itself as the mechanism for accessing the Azure Databricks data. The AWS-side engine still needs a supported way to consume the shared data.&lt;/P&gt;&lt;P&gt;I would therefore validate which AWS compute engine needs to query those Glue tables first, and then determine whether it can consume the OpenSharing share directly or requires an integration layer.&lt;/P&gt;&lt;P&gt;Also consider cross-cloud network egress cost and network controls, even though the architecture avoids maintaining another copy of the dataset.&lt;/P&gt;</description>
      <pubDate>Wed, 07 Oct 2026 18:18:33 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172195#M56593</guid>
      <dc:creator>vs4</dc:creator>
      <dc:date>2026-10-07T18:18:33Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172209#M56598</link>
      <description>&lt;DIV&gt;Thank you so much for all your help and guidance. I truly appreciate your responses and insights. I’ll definitely keep this in mind and follow up with the source team first.&lt;/DIV&gt;</description>
      <pubDate>Wed, 07 Oct 2026 20:48:44 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172209#M56598</guid>
      <dc:creator>priya9896</dc:creator>
      <dc:date>2026-10-07T20:48:44Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172268#M56609</link>
      <description>&lt;P&gt;Thanks for the detailed explanation Louis. this really helps in understanding the concept more clearly. The distinction between “no replication” and “no data movement” is particularly important here.&lt;/P&gt;&lt;P&gt;If OpenSharing is used with AWS Glue, the data can remain in Azure while the AWS workload reads it on demand. One thing I would be interested in understanding is how this behaves for larger analytical workloads.&lt;/P&gt;&lt;P&gt;For example, if the AWS consumers frequently query large Delta tables, would it be better to expose curated/filtered shared views or partition-filtered tables through OpenSharing rather than allowing consumers to query the full datasets? That seems like it could help reduce cross-cloud data transfer, latency, and egress costs.&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 09:45:51 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172268#M56609</guid>
      <dc:creator>Aravind_Reddy</dc:creator>
      <dc:date>2026-10-08T09:45:51Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172280#M56612</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/192529"&gt;@Aravind_Reddy&lt;/a&gt;,&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Short answer: yes, share the smallest useful slice, but the mechanism matters, because OpenSharing serves tables and views very differently to an open recipient like Glue. And as &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/149393"&gt;@vs4&lt;/a&gt; said above, settle which AWS engine is reading before you design the share.&lt;/P&gt;
&lt;P&gt;Order of preference:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Partition-limited tables, when the access pattern has stable boundaries (date, region, tenant). The share carries a partition spec, static (&lt;CODE&gt;PARTITION (year = '2026')&lt;/CODE&gt;) or recipient-driven (&lt;CODE&gt;PARTITION (region = CURRENT_RECIPIENT().region)&lt;/CODE&gt;), and the server never lists files outside it. No provider compute, nothing for the consumer to get wrong. The consumer's own filters also travel to the server as predicate hints: partition columns prune well, other columns are best effort and file-level. Caveats: the table needs Hive-style partitions (you can't partition-filter a liquid-clustered table), a filter disqualifies the table from directory-based access (below), and a big table means a big file list per query with limits on file counts, so keep it compacted.&lt;/LI&gt;
&lt;LI&gt;A curated Delta table, when the subset is more than a partition predicate. The source team maintains a narrow table in Azure (needed columns, rows, retention window) and shares that. It gets everything in point 1 with no per-query compute on their side.&lt;/LI&gt;
&lt;LI&gt;A materialized view, for a stable reporting shape that's queried often. The heavy query runs on the refresh schedule instead of on every read; open recipients get the current snapshot.&lt;/LI&gt;
&lt;LI&gt;Views, for governance (row filters or masking by recipient via &lt;CODE&gt;CURRENT_RECIPIENT()&lt;/CODE&gt;) or for a slice the source team won't materialize. Don't assume a view is faster. For a Databricks-to-Open recipient, every query runs on the provider's serverless compute, the result is fully materialized regardless of the consumer's filters (no predicate or &lt;CODE&gt;LIMIT&lt;/CODE&gt; pushdown), parked temporarily in the provider's storage and served through short-lived URLs. The provider pays, and can see what each recipient costs by joining &lt;CODE&gt;system.sharing.materialization_history&lt;/CODE&gt; to &lt;CODE&gt;system.billing.usage&lt;/CODE&gt;. A view as wide as the table is strictly worse than the table, plain and simple, and "frequently queried" means "frequently materialized."&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Two things apply regardless. Consumer-side hygiene: select explicit columns (the connector fetches Parquet column chunks over range requests, so fewer columns means fewer bytes) and filter on partition columns that match the workload; a partition boundary the queries don't use buys you nothing. And the newer directory-based access mode: an unfiltered Delta table shared &lt;CODE&gt;WITH HISTORY&lt;/CODE&gt; can be served with a short-lived, path-scoped credential, so the recipient's own Delta reader does the planning and data skipping, liquid clustering included. Check that your Glue connector version supports it (recent protocol addition, URL mode is the fallback), and know that the credential covers the Delta log, which includes commit history and deleted data not yet vacuumed.&lt;/P&gt;
&lt;P&gt;For the POC, run the same representative queries three ways: unrestricted table, partition-limited share, curated or aggregated table. Measure bytes transferred, end-to-end latency, Glue runtime, provider compute (the materialization table above), and egress. If the workload keeps scanning most of the source data, cross-cloud reads stay expensive however you slice it; the honest fixes are upstream curation, compute in Azure, or the replica the source team has already ruled out.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Create shares (partition specs, recipient-property filtering, views, cloud token eligibility): &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/create-share" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/create-share&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;OpenSharing overview, including who pays for view materialization and table limitations: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Materialization history system table: &lt;A href="https://learn.microsoft.com/azure/databricks/admin/system-tables/materialization" target="_blank"&gt;https://learn.microsoft.com/azure/databricks/admin/system-tables/materialization&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Monitor and manage OpenSharing egress costs: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/opensharing/manage-egress&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Delta Sharing protocol (predicate hints and access modes): &lt;A href="https://github.com/delta-io/delta-sharing/blob/main/PROTOCOL.md" target="_blank"&gt;https://github.com/delta-io/delta-sharing/blob/main/PROTOCOL.md&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 12:24:56 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172280#M56612</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-08T12:24:56Z</dc:date>
    </item>
    <item>
      <title>Re: Azure to AWS</title>
      <link>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172281#M56613</link>
      <description>&lt;P&gt;Also, if you feel your question has been answered please "Accept as Solution" so that others can benefit.&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Cheers, Lou.&lt;/P&gt;</description>
      <pubDate>Thu, 08 Oct 2026 12:25:38 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/azure-to-aws/m-p/172281#M56613</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-08T12:25:38Z</dc:date>
    </item>
  </channel>
</rss>

