- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
Monday
Hello @SKiruthiga, I took a look at both internal and external documentation and here is what I found.
Short version: the Snowflake Connector for Spark does support workload identity federation, starting with version 3.1.7 (certified against JDBC 3.28.0). The harder question is whether your Azure Databricks cluster can hand the connector a usable Azure identity. Those are two different things.
The connector options look like this, and the same options work for df.write:
sfOptions = {
"sfURL": "<account_identifier>.snowflakecomputing.com",
"sfUser": "<snowflake_service_user>",
"sfAuthenticator": "WORKLOAD_IDENTITY",
"sfWorkloadIdentityProvider": "AZURE",
"sfDatabase": "<database>",
"sfSchema": "<schema>",
"sfWarehouse": "<warehouse>"
}
df = spark.read.format("snowflake").options(**sfOptions).option("dbtable", "<table>").load()
On the Snowflake side, create a service user whose identity matches the Azure identity the job runs as, then grant it only the role and privileges it needs:
CREATE USER <snowflake_service_user>
TYPE = SERVICE
WORKLOAD_IDENTITY = (
TYPE = AZURE
ISSUER = 'https://login.microsoftonline.com/<tenant_id>/v2.0'
SUBJECT = '<managed_identity_object_id>'
);
Now a caveat. AZURE mode expects to fetch a token from the Azure Instance Metadata Service for a managed identity attached to the VM. On Databricks classic compute those VMs sit in the Databricks-managed resource group, so you can't attach your own identity to them, and serverless has no VM you control at all. I don't see anything public that guarantees this endpoint is available to the connector, and I'd expect it to fail. If it does, tweaking the two Spark options won't fix it.
The route I'd try instead is OIDC mode with a Unity Catalog service credential supplying the token. A service credential wraps a managed identity (through an Access Connector for Azure Databricks) and gives your code a TokenCredential. Rough shape:
- Tenant admin, one time: consent to Snowflake's multi-tenant Entra app (link in the Snowflake WIF doc). That app is the audience Snowflake expects on Azure tokens.
- Azure: create an Access Connector for Azure Databricks and note its managed identity's Object ID.
- Databricks: create a UC service credential on that access connector and grant ACCESS to the job's principal.
- In the job, request a token and pass it through:
cred = dbutils.credentials.getServiceCredentialsProvider("snowflake-wif")
token = cred.get_token("api://fd3f753b-eed3-462c-b6a7-a4b5bb650aad/.default").token
sfOptions["sfWorkloadIdentityProvider"] = "OIDC"
sfOptions["sfToken"] = token
- Snowflake: decode one token first (the WIF doc has a jq one-liner) and copy the exact
issandsubclaims intoISSUERandSUBJECT. Managed identity tokens sometimes carry the v1 issuer (sts.windows.net/...) rather than v2, and Snowflake matches exactly. If you useTYPE = OIDCon the user instead ofTYPE = AZURE, setOIDC_AUDIENCE_LISTto the token's audience too.
I haven't run the OIDC plus service credential combination end to end, so treat it as the approach I'd try rather than a confirmed recipe. And check which connector version your Databricks Runtime bundles (ls /databricks/jars | grep -i snowflake on the cluster). If it's older than 3.1.7 you'll need to attach the newer connector and JDBC jars and confirm they win over the bundled ones.
A few smaller notes. Entra tokens last about an hour; that's fine for a batch job that fetches a fresh one per run, but a long-running streaming job needs a refresh plan. Don't confuse Snowflake WIF with Databricks OAuth token federation, which is the reverse direction (external workloads authenticating into Databricks) and a separate trust setup. And whichever path you pick, validate with a SELECT 1 read and a tiny write before migrating the real job.
If WIF is a dead end in your setup, Snowflake External OAuth with a short-lived Entra token (sfAuthenticator = oauth) is the fallback. It needs a Snowflake security integration and its own token refresh handling, but it still gets you off stored passwords. If your need is mostly reads, Lakehouse Federation's Snowflake connection with OAuth keeps auth out of your job code entirely.
References:
- Snowflake WIF overview, including Azure setup and supported drivers: https://docs.snowflake.com/en/user-guide/workload-identity-federation
- Snowflake Spark connector options: https://docs.snowflake.com/en/user-guide/spark-connector-use
- Databricks: read and write Snowflake data: https://learn.microsoft.com/en-us/azure/databricks/connect/external-systems/snowflake
- Create UC service credentials on Azure: https://learn.microsoft.com/en-us/azure/databricks/connect/unity-catalog/cloud-services/service-cred...
- Use service credentials from code: https://learn.microsoft.com/en-us/azure/databricks/connect/unity-catalog/cloud-services/use-service-...
If you get it working, please post back with the ISSUER and SUBJECT shape you ended up with. It'd save the next person a lot of trial and error.
Regards, Louis.