Ashwin_DSA
Databricks Employee
Databricks Employee

Hi @ivanvyd,

Thanks for the question. The short version is that filtering by current_user() usually keeps the broadcast small, but the thing to factor in the design is the lookup in the entitlement table, not the broadcast itself. current_user() resolves to a single constant for the query, so the subquery runs once, not per row. It returns only that user's handful of keys, and that small result is broadcast to filter the fact table. So even with millions of mappings, the broadcast side stays tiny as long as each user maps to only a few keys.

The expensive part is reading the millions-row entitlement table to find those few rows on every query. That is where I think the layout matters. Cluster the entitlement table on user_email, liquid clustering is the easy choice, so Delta can skip to the user's files instead of scanning the whole table. Keep it compacted with OPTIMIZE, avoid a pile of small files, and make sure you are not wrapping user_email in a function or cast that defeats data skipping. With that in place, the per-query lookup prunes to a few files, and the pattern should hold at scale.

If you hit a case where per-user sets are legitimately large, that is the point to rethink the layout, for example, scoping by a group or an attribute rather than an explicit per-row mapping. But for a few mappings per user, a well-clustered control table is the right call and should scale fine.

 

Regards,
Ashwin | Delivery Solution Architect @ Databricks
Helping you build and scale the Data Intelligence Platform.
***Opinions are my own***