cancel
Showing results for 
Search instead for 
Did you mean: 
Data Governance
Join discussions on data governance practices, compliance, and security within the Databricks Community. Exchange strategies and insights to ensure data integrity and regulatory compliance.
cancel
Showing results for 
Search instead for 
Did you mean: 

How to override ABAC policies

cpayne_vax
Contributor

I'm struggling with the best way to handle ABAC inheritance. Here's the situation:

We have tags deployed in table columns across a dozen catalogs and their schemas. In general, certain tags should have column masking policies applied so that our user base cannot see unmasked data. However, there are overrides. 

Imagine:
Catalog A
-> Schema 1
-> Schema 2
Catalog B
-> Schema 3

I have a group "Unmasked users" that get unmasking everywhere, so I've applied policies on both catalogs with "Unmasked users" in the EXCEPT list. But for Schema 1 specifically, I also need "Privileged Users" to have UNMASKING. 

So I create a policy to apply at the Schema 1 level with "Privileged Users" in the EXCEPT list. But here's the catch: because "Schema 1" inherits the policy that is applied on Catalog A which only allows "Unmasked users," then the more restrictive policy applies and "Privileged Users" can't see masked data.

The ways I see to get what I'm after:

  1. Stop applying policies at the catalog level. This would make management immensely more complicated.
  2. Add "Privileged Users" to the policy at Catalog A. This would mean this group also gets unmasked access to Schema 2, which is unacceptable.

Ideally I'd be able to override the inherited policy, but I see no way to do this. What's the best way to accomplish this goal? (And bear in mind that my example here is very simple, but I have over a dozen catalogs with hundreds of schemas, which makes option 1 unbearable.)

1 ACCEPTED SOLUTION

Accepted Solutions

ShamenParis
Contributor III

@cpayne_vax you've mapped this out perfectly.

@ThomazNeto really nailed the WHEN NOT approach, which is the real lifesaver here. By using WHEN NOT has_tag('mask_scope') for your default policy, you don't even need to manually apply a 'standard' tag to Catalogs A, B, and C. The default just applies everywhere automatically, and you only lift a finger to tag your exceptions (Schemas 1, 2, and 4).

Regarding your Finance,Security snag: your workaround is exactly the right approach. If you tried applying two separate tags to trigger two separate policies on the same column, Databricks wouldn't grant access to both—it would throw a MULTIPLE_MASKS error and lock it down. Creating a combined tag value and a dedicated policy for that specific combo is the native way to handle it.

As an architect, my only tweak to your sketch is how you deploy it: deploy all those policy permutations to all your catalogs.

Don't just apply the Finance policy to Catalog A and the Combo policy to Catalog B. Apply them all everywhere. That way, if a schema in Catalog C suddenly needs Finance,Security access tomorrow, you just apply the tag and the logic is already sitting there waiting for it.

Manage those few combination policies centrally (via Terraform or a Python loop) and you're good to go. You've got this entirely straight in your head!

View solution in original post

6 REPLIES 6

ShamenParis
Contributor III

Hi @cpayne_vax ,

I’ve run into this exact inheritance trap while architecting platforms with Unity Catalog.

Because UC always prioritizes the most restrictive policy to prevent accidental data leaks, fighting the catalog-to-schema inheritance is a losing battle. If Catalog A enforces a policy based on a tag, Schema 1 cannot safely override it if both rules trigger off the exact same tag condition.

Here is the cleanest way I handle this without creating a management nightmare: Stop trying to override the policy, and instead branch the tags.

Keep all your policy management at the Catalog level, but use mutually exclusive tag values to trigger them using the WHEN clause.

1. Define Mutually Exclusive Tags

Instead of a single boolean tag, use a key-value tag structure to represent the access tiers:

  • Tag Key: Security_Tier -> Value: Standard

  • Tag Key: Security_Tier -> Value: Privileged_Exception

2. Create Tag-Conditioned Policies at the Catalog Level

Create two distinct masking policies directly on the Catalog, using WHEN has_tag_value() to ensure they never overlap.

The Standard Policy:

CREATE OR REPLACE POLICY catalog_standard_masking
ON CATALOG `catalog_a`
COLUMN MASK default_masking_udf
TO `account users`
EXCEPT `Unmasked users`
FOR COLUMNS
WHEN has_tag_value('Security_Tier', 'Standard');

The Privileged Policy:

CREATE OR REPLACE POLICY catalog_privileged_masking
ON CATALOG `catalog_a`
COLUMN MASK privileged_masking_udf
TO `account users`
EXCEPT `Unmasked users`, `Privileged Users`
FOR COLUMNS
WHEN has_tag_value('Security_Tier', 'Privileged_Exception');

3. Apply the Tags Downstream

Apply Security_Tier: Standard to your standard schemas (like Schema 2) and apply Security_Tier: Privileged_Exception to Schema 1.

By doing this, you avoid the inheritance clash entirely. A column in Schema 1 only triggers the privileged policy, so the restrictive standard policy never fires. You manage just two policies at the very top of your dozen catalogs, and you control all your exceptions purely by tagging the data correctly downstream.

ThomazNeto
Databricks Partner

Hi,

You're reading the behavior right, and there is no override mechanism. The docs are explicit about how in-scope policies combine: the engine "Identifies all policies whose scope covers the queried table", checks TO/EXCEPT per policy, and then "Only one distinct row filter can resolve at runtime for a given table and a given user, and only one distinct column mask can resolve for a given column and a given user." Nothing at a child level cancels a parent policy. In your case the catalog policy simply still applies to Privileged Users (they're in its TO, not its EXCEPT), so they get masked. If two different masks ever did resolve for the same user, you wouldn't get "most restrictive wins", you'd get an error (MULTIPLE_MASKS) and the table becomes unreadable for that user.
link 1
link 2

The documented lever you're missing is the WHEN clause combined with tag inheritance. has_tag in WHEN "checks tags set directly on the table or inherited from a parent catalog or schema", and NOT is allowed in the condition. So you keep one policy per catalog and carve schemas out of it with a governed tag:

  1. Create a governed tag, say mask_scope, and set mask_scope = privileged on Schema 1. Every table in it inherits the tag for policy evaluation.
  2. Catalog A policy: ... TO account users EXCEPT unmasked_users FOR TABLES WHEN NOT has_tag_value('mask_scope', 'privileged') MATCH COLUMNS has_tag('pii') AS c ON COLUMN c
  3. Schema 1 policy: same mask UDF, TO account users EXCEPT unmasked_users, privileged_users FOR TABLES MATCH COLUMNS has_tag('pii') AS c ON COLUMN c

Schema 2 keeps the catalog policy untouched, Schema 1 is governed only by its own policy, and since both policies use the same UDF there's no conflict risk even during the transition. SHOW EFFECTIVE POLICIES ON TABLE ... is how you verify what actually lands on a table.
link 3
link 4

Three caveats from the docs worth respecting at your scale. Tagging is a security boundary ("If a user can change tags on a data asset, they can change which policies apply to it"), so lock ASSIGN on that tag to your governance group. Tag changes take a few minutes to propagate. And there's a quota of 20 principals per policy across TO and EXCEPT, so keep the exemptions as groups, not individual users.

With a dozen catalogs and hundreds of schemas, you end up with one policy per catalog plus one per exception schema, which is about as small as this can get.

Thomaz A. Rossito Neto
Principal Data Architect & AI Strategy — CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto

cpayne_vax
Contributor

Thank you both, @ShamenParis and @ThomazNeto. This really helps. Let me restate and generalize what I think I'm hearing so that I get it straight in my head.

The reality is that I don't have just 1 "privileged users" group. I have many, each allowed to view unmasked data related to their job functions. Imagine groups for: Finance team, Patient team, Security team, etc. So I create a managed tag called "mask_scope" and each schema will have a value for this tag that corresponds to the team(s) that can see unmasked data. Some schemas will have multiple, and this presents a snag. Tags can't have multiple values when applied, which means I'd have to create a CROSS apply of sorts, where allowable values could be "Finance,Security" and "Patients,Security" and "Finance."

So all that said, here's my new sketch:

Catalog A
-> Schema 1
-> Schema 2
-> Schema 3
Catalog B
-> Schema 4
-> Schema 5
Catalog C
-> Schema 6

Requirements:
  • Finance can see Schema 1
  • Security can see Schema 2
  • Finance and Security can see Schema 4
  • No one can read A.3, B.5, or C.6
  • Normies can't read anything

Tags:

  • Catalog A,B,C: standard
  • Schema 1: mask_scope = Finance
  • Schema 2: mask_scope = Security
  • Schema 4: mask_scope = Finance,Security

Policies:

  1. Applied on catalogs A,B,C to account users, WHEN tag = standard
  2. Applied on catalog A to account users, EXCEPT Finance, WHEN tag = Finance
  3. Applied on catalog A to account users, EXCEPT Security, WHEN tag = Security
  4. Applied on catalog B to account users, EXCEPT Finance and Security, WHEN tag = Finance,Security

Effective:

  • Schemas 3, 5, 6 inherit standard tag, no one can read them
  • Finance can read Schema 1 because of policy 2
  • Security can read Schema 2 because of policy 3
  • Finance and Security can read schema 4 because of policy 4

Am I missing anything? Again, thank you both for helping expand my thinking on this!

ShamenParis
Contributor III

@cpayne_vax you've mapped this out perfectly.

@ThomazNeto really nailed the WHEN NOT approach, which is the real lifesaver here. By using WHEN NOT has_tag('mask_scope') for your default policy, you don't even need to manually apply a 'standard' tag to Catalogs A, B, and C. The default just applies everywhere automatically, and you only lift a finger to tag your exceptions (Schemas 1, 2, and 4).

Regarding your Finance,Security snag: your workaround is exactly the right approach. If you tried applying two separate tags to trigger two separate policies on the same column, Databricks wouldn't grant access to both—it would throw a MULTIPLE_MASKS error and lock it down. Creating a combined tag value and a dedicated policy for that specific combo is the native way to handle it.

As an architect, my only tweak to your sketch is how you deploy it: deploy all those policy permutations to all your catalogs.

Don't just apply the Finance policy to Catalog A and the Combo policy to Catalog B. Apply them all everywhere. That way, if a schema in Catalog C suddenly needs Finance,Security access tomorrow, you just apply the tag and the logic is already sitting there waiting for it.

Manage those few combination policies centrally (via Terraform or a Python loop) and you're good to go. You've got this entirely straight in your head!

cpayne_vax
Contributor

Awesome. Thank you both again, this is fairly simple to implement with our terraform code. You both gave great answers so I don't know who to accept as solution! I'll just pick the last response.

Thanks again!

Srini_Pesala
Databricks Partner

The scalable way to handle this is to keep the catalog-level policy, but use a governed tag and a WHEN condition to carve out the schema-level exception.

You don't need to remove the catalog-level policy or add "Privileged Users" to the catalog-level EXCEPT list.

For example, define a governed tag such as:

mask_scope = privileged

Apply this tag to Schema 1.

Because tags are inherited for ABAC policy evaluation, the tables under Schema 1 will effectively have this tag as well.

Then modify the Catalog A policy so that it does not apply to objects with this exception tag.

Conceptually:

Catalog A policy:

TO: Users
EXCEPT: Unmasked Users

WHEN:

NOT has_tag_value('mask_scope', 'privileged')

MATCH COLUMNS:

has_tag('sensitive')

This means the catalog-level policy continues to protect the entire catalog, except for objects under schemas that have mask_scope = privileged.

Then create a Schema 1 policy for the specific exception:

TO: Users
EXCEPT: Unmasked Users, Privileged Users

MATCH COLUMNS:

has_tag('sensitive')

The result is:

Catalog A
Schema 1
Unmasked Users → Unmasked
Privileged Users → Unmasked
Other Users → Masked

Schema 2
Unmasked Users → Unmasked
Privileged Users → Masked
Other Users → Masked

This gives you the desired behavior without having to manage policies individually for hundreds of schemas.

The important idea is that the lower-level policy isn't actually "overriding" the inherited policy. Instead, the inherited catalog policy is conditionally prevented from applying to the exception schema, and the schema-level policy handles that schema.

This also scales well because the default behavior remains at the catalog level. You only need to tag the schemas that require an exception.

For example:

Catalog A
├── Schema 1 → mask_scope = privileged
├── Schema 2 → no exception tag
├── Schema 3 → no exception tag
└── Schema 4 → no exception tag

Only Schema 1 requires additional configuration.

I would also recommend using:

SHOW EFFECTIVE POLICIES ON SCHEMA Catalog A.Schema 1;

to verify which policies are being inherited and applied.

This approach aligns with the Databricks recommendation to keep policies at the highest applicable scope while using governed tags and policy conditions to handle exceptions. Governed tags can be inherited from catalogs and schemas during ABAC evaluation.

So, in short:

Keep the catalog-level policy → Add an exception tag to Schema 1 → Use WHEN NOT on the catalog policy → Add the schema-specific policy for Privileged Users.

This avoids the administrative overhead of moving everything to schema-level policies and prevents Privileged Users from getting unmasked access to Schema 2.