Louis_Frolio
Databricks Employee
Databricks Employee

Hello @priya9896, I took a look at both internal and external documentation and here is what I found.

Straight answer: don't plan on the Glue Data Catalog as a federation layer over Azure Databricks. Glue's Delta integration and the Glue Catalog itself are built around tables in S3. Even Glue catalog federation to Unity Catalog's Iceberg REST endpoint works by having Lake Formation vend scoped credentials to data stored in S3, and AWS's walkthrough for it lists a Databricks workspace on AWS as a prerequisite. So: can a Glue Spark job read Azure Databricks data without copying it? Yes. Can you get Glue Catalog tables that Athena queries without a copy? No. Your two goals pull against each other, and one has to give.

With "no copy" as the hard constraint, the supported pattern is Delta Sharing, which Databricks renamed OpenSharing in June 2026. Same protocol; the package is still delta-sharing-spark and the format is still deltaSharing.

If your AWS consumers have a Unity Catalog-enabled Databricks workspace, use Databricks-to-Databricks sharing. The shared tables show up read-only in their catalog and there's no credential file to manage.

If Glue is the actual consumer, use Databricks-to-Open sharing. The source team creates a share with only the tables you need (partition filters or shared views if they want to narrow rows or columns), creates a recipient (bearer token or OIDC), and sends you a credential file. A Glue Spark job then reads it like this, with the credential file locked down since it is a bearer token:

df = spark.read.format("deltaSharing").load("s3://<secure-bucket>/config.share#<share>.<schema>.<table>")

The same credential file carries an icebergEndpoint, so Iceberg clients (Spark with an Iceberg REST catalog, PyIceberg, Trino, Snowflake) can read the share too.

Either way, Databricks returns the table's storage location with temporary cloud credentials, and your compute reads directly from Azure storage. Nothing persists in AWS. That's no replication, but it isn't zero network movement: rows cross the cloud boundary on every query. Prototype it from a Glue job, measure latency and throughput, and settle egress with the source team up front, because your cloud vendor may charge egress fees for sharing across clouds, and those land on whoever owns the storage, which is them. Databricks documents Cloudflare R2 as an egress-free option (a replica, so their call), and with SecureConnect, Databricks bills the data transfer rather than the cloud vendor; it's in Public Preview, so ask about it. Also pin down what "no ingestion into AWS" means. If it means no persisted copy, OpenSharing fits. If it means no bytes leave Azure, nothing here works and the consumers need to compute in Azure.

On the other patterns:

  • Unity Catalog's Iceberg REST catalog or Unity REST API with a service principal. Works from Spark, but the source team has to enable external data access on their metastore and manage a principal for you. OpenSharing is the purpose-built version for a separate consuming team: scoped to a share, revocable, audited.
  • JDBC to a Databricks SQL warehouse. Fine for targeted extracts, but every read runs on the source team's warehouse, rows move over JDBC rather than parallel Parquet reads, and the result still isn't a native Glue Catalog table Athena can query.
  • Direct storage access. AWS does publish an Azure Data Lake Storage connector for Glue that can read Delta tables straight from ADLS Gen2, but it runs on storage credentials, not Unity Catalog permissions, so it bypasses the source team's governance entirely. Most source teams won't agree, and I wouldn't ask.
  • Lakehouse Federation runs the other direction (Databricks querying external systems), so it doesn't help here.

If your consumers truly need glueContext.create_data_frame.from_catalog(...) or Athena, that's a different problem, and the only route is a governed, scheduled replica into S3 that Glue then catalogs. That's replication, which the source team has ruled out, so I'd have that conversation with them rather than engineer around it. First, though, confirm whether the consumers need Glue Catalog metadata or just governed read access. In my experience it's usually the latter, and OpenSharing covers it.

References:

Regards, Louis.