cancel
Showing results forย 
Search instead forย 
Did you mean:ย 
Data Governance
Join discussions on data governance practices, compliance, and security within the Databricks Community. Exchange strategies and insights to ensure data integrity and regulatory compliance.
cancel
Showing results forย 
Search instead forย 
Did you mean:ย 

Personal Compute Policies

Ankitkalra40
New Contributor III

I am working with a client to set up their budget policies and compute governance.

They have a team of Engineers and Scientists that work across different versions of various libraries like python, pandas, numpy, scikit, mlflow etc.

Since each user requires diff versions, having one init script defeats the purpose. Also, any library and the version to be installed is supposed to be validated by InfoSec. 

I have proposed a rather niche solution to this that utilizes DBX app as the interface and syncs the approved libraries via Lakebase instance.

We are testing this out but wanted to check if there is something easier available within Databricks at a platform level.

Thanks in advance!

4 REPLIES 4

ThomazNeto
Databricks Partner

Hi Ankit,

Before building anything custom, have a look at three things that already exist in the platform. Together they cover most of what you described.

  1. Workspace base environments. A workspace admin defines a YAML with the serverless environment version and a pinned list of Python dependencies (requirements file, wheels in a UC volume, an internal index-url). Databricks pre-builds and caches it, users pick it from the Environment panel, and admins can star one as the workspace default. You're limited to 10 per workspace, so think "approved profiles" (data-eng, ml-sklearn, ml-mlflow) rather than one per person. Base environments also work on classic compute through "manage dependencies using environments". This is probably the closest thing to what you're prototyping, without the app and the Lakebase sync.
    https://docs.databricks.com/aws/en/admin/workspace-settings/base-environment
    https://docs.databricks.com/aws/en/compute/serverless/dependencies

  2. Compute policies with libraries. For classic compute, a policy can carry up to 500 libraries (PyPI, wheels in volumes, requirements.txt) that get installed automatically, and the docs are explicit about the side effect you want: "Users can't install or uninstall compute-scoped libraries on compute that use this policy." Databricks also recommends policies over init scripts for library installs. Note this blocks compute-scoped libraries; notebook-scoped %pip still works unless you cut off the index (see next point).
    https://docs.databricks.com/aws/en/admin/clusters/policies

  3. Control the source, not the list. The docs say private mirrors like Nexus or Artifactory are supported via --index-url in %pip and in base environment YAML, and admins can set a private repo as the default pip source for serverless. If InfoSec curates what's in the mirror, the approval happens once, at the repository, and every install path inherits it. That's cleaner than approving versions inside Databricks.
    https://docs.databricks.com/aws/en/libraries/

One more, since you tagged Unity Catalog: on standard access mode, JARs, Maven coordinates and init scripts need to be in the UC artifact allowlist (MANAGE ALLOWLIST privilege). It doesn't cover PyPI, so for Python the mirror is the control point.
https://docs.databricks.com/aws/en/data-governance/unity-catalog/manage-privileges/allowlist

Your app might still make sense as the request/approval UI, but I'd have it write base environments and policy libraries rather than manage installs itself.

Hope this helps.

Thomaz A. Rossito Neto
Principal Data Architect & AI Strategy โ€” CI&T
thomazn@ciandt.com
linkedin.com/in/thomaz-antonio-rossito-neto

Ankitkalra40
New Contributor III

Hi Thomaz, 
I agree with most of your recommendations which are in line with the best practices as well.
However, the client is a bit strict about their data policies and Serverless is disabled at their org level.

This is the primary reason a lot of above can not be directly implemented or tested for their requirement.

Is there any other alternative that you can help with? Happy to connect separately on this as well

balajij8
Esteemed Contributor II

@Ankitkalra40 

You can handle it natively using Compute Policies with Enforced Libraries instead of building a custom DBX app to sync libraries. Workspace admins can build per team compute policies that enforce specific library versions. The key is that users on these policies are completely locked out of installing or uninstalling libraries on this compute. More details here

You set it up directly in the policy's Libraries tab that accepts up to 500 libraries per policy. Instead of using a monolithic init script, you can create dedicated policies per team / dependency matrix eg., one policy mapped to Engineers and another for Scientists

You can pair these policies with Unity Catalog Volumes to address the InfoSec approval process. You can store the InfoSec-validated .whl, .jar and requirements.txt files in a UC Volume and use Unity Catalog grants to restrict READ access to the appropriate teams. More details here

Compute policies can point directly to these volume paths instead of public repositories. It creates a tightly gated, native workflow - InfoSec reviews a library and drops the package into the secure volume, the team's compute policy picks it up and the engineers cannot deviate.

ivanvyd
New Contributor III

@Ankitkalra40 since Serverless is disabled, there is one newer native option worth checking: as of the September 17 release, base environments are supported on classic compute in Beta.

I haven't tried it yet, honestly, but as you might want to consider it as an alternative to other answers.

On DBR 19+ with Standard access mode, you can use Dependency mode = Environments and provide admin-managed workspace base environments.

One caveat for the InfoSec requirement: this is not a hard allowlist by itself, because users can still add notebook-level dependencies. If only approved packages must be installable, I'd keep the trust boundary at the package source: enforced compute-policy libraries / approved wheels in UC Volumes, plus a curated private package repository and egress restrictions.

Also worth noting that Environments mode currently targets interactive all-purpose notebooks; classic jobs don't use it yet.

Ivan Vydrin
Lead Software & AI Engineer ยท Tech Fabric LLC