Tuesday
Hi everyone,
We're evaluating an architecture pattern and would appreciate any guidance or recommendations.
Has anyone implemented a similar cross-cloud pattern? Specifically:
Thanks in advance!
Tuesday
Hi everyone,
We're evaluating an architecture pattern and would appreciate any guidance or recommendations.
Current state:
Source data resides in Azure Databricks.
Consumers are in AWS .
The source team allows read-only access but does not allow data replication or ingestion into AWS.
Goal: We would like AWS consumers to access the data through Glue Catalog tables while keeping the data in Azure Databricks.
Has anyone implemented a similar cross-cloud pattern? Specifically:
Can AWS Glue be used to access Azure Databricks data without copying it?
Are there recommended approaches such as federation, Delta Sharing, custom connectors, JDBC/SQL endpoints, or other patterns?
Appreciate all your time and inputs regarding . Thanks in advance!
yesterday
Great explanation! I'd add one important consideration from a cross-cloud architecture and data governance perspective.
Before implementing OpenSharing with AWS Glue, I would clarify exactly what the source team means by "no replication or ingestion into AWS."
There is an important difference between avoiding a permanent replica in S3 and prohibiting any temporary data persistence or processing within AWS.
Even when using OpenSharing, the data still crosses cloud boundaries. Depending on the Glue job configuration, intermediate Spark data could also be written to temporary storage.
I'd suggest validating three things during a small proof of concept:
Governance: Confirm whether temporary processing and potential data spilling in AWS are permitted.
Performance: Measure cross-cloud transfer costs, query latency and the volume of data transferred.
Operations: Validate credentials, network connectivity and how shared-table schema changes affect downstream consumers.
If the requirement is strictly read-only access without a permanent replica, OpenSharing is worth evaluating.
However, if no data is permitted to leave Azure, I would consider keeping the processing in Azure and exposing only approved results.
One question: Is the Glue Catalog requirement driven by Athena integration, or could the AWS consumers work directly with shared datasets through Glue Spark?
yesterday
Hello @priya9896, I took a look at both internal and external documentation and here is what I found.
Straight answer: don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itself are built around tables in S3. Even Glue catalog federation to Unity Catalog's Iceberg REST endpoint works by having Lake Formation vend scoped credentials to data stored in S3, and AWS's walkthrough for it lists a Databricks workspace on AWS as a prerequisite. So: can a Glue Spark job read Azure Databricks data without copying it? Yes. Can you get Glue Catalog tables that Athena queries without a copy? No. Your two goals pull against each other, and one has to give.
With "no copy" as the hard constraint, the supported pattern is Delta Sharing, which Databricks renamed OpenSharing in June 2026. Same protocol; the package is still delta-sharing-spark and the format is still deltaSharing.
If your AWS consumers have a Unity Catalog-enabled Databricks workspace, use Databricks-to-Databricks sharing. The shared tables show up read-only in their catalog and there's no credential file to manage.
If Glue is the actual consumer, use Databricks-to-Open sharing. The source team creates a share with only the tables you need (partition filters or shared views if they want to narrow rows or columns), creates a recipient (bearer token or OIDC), and sends you a credential file. A Glue Spark job then reads it like this, with the credential file locked down since it is a bearer token:
df = spark.read.format("deltaSharing").load("s3://<secure-bucket>/config.share#<share>.<schema>.<table>")
The same credential file carries an icebergEndpoint, so Iceberg clients (Spark with an Iceberg REST catalog, PyIceberg, Trino, Snowflake) can read the share too.
Either way, Databricks returns the table's storage location with temporary cloud credentials, and your compute reads directly from Azure storage. Nothing persists in AWS. That's no replication, but it isn't zero network movement: rows cross the cloud boundary on every query. Prototype it from a Glue job, measure latency and throughput, and settle egress with the source team up front, because your cloud vendor may charge egress fees for sharing across clouds, and those land on whoever owns the storage, which is them. Databricks documents Cloudflare R2 as an egress-free option (a replica, so their call), and with SecureConnect, Databricks bills the data transfer rather than the cloud vendor; it's in Public Preview, so ask about it. Also pin down what "no ingestion into AWS" means. If it means no persisted copy, OpenSharing fits. If it means no bytes leave Azure, nothing here works and the consumers need to compute in Azure.
On the other patterns:
If your consumers truly need glueContext.create_data_frame.from_catalog(...) or Athena, that's a different problem, and the only route is a governed, scheduled replica into S3 that Glue then catalogs. That's replication, which the source team has ruled out, so I'd have that conversation with them rather than engineer around it. First, though, confirm whether the consumers need Glue Catalog metadata or just governed read access. In my experience it's usually the latter, and OpenSharing covers it.
References:
Regards, Louis.