cancel
Showing results forย 
Search instead forย 
Did you mean:ย 
Community Articles
Dive into a collaborative space where members like YOU can exchange knowledge, tips, and best practices. Join the conversation today and unlock a wealth of collective wisdom to enhance your experience and drive success.
cancel
Showing results forย 
Search instead forย 
Did you mean:ย 

Questions for Choosing a Databricks Partner, and what I'd add after living with the platform

ericka-lorenz
New Contributor II

I came across this checklist for hiring a Databricks implementation partner, and it turned out to be the best in-house build checklist I've read in a while. 27 Questions to Choose a Databricks Implementation Partner looks like procurement filler at first, but almost every question is really about the implementation. So if your own team is the one building, and a lot of us here are the in-house team rather than the buyer, the list points straight at you.

The reason it holds up is that a bad implementation rarely hurts you in the day rate. It hurts months later. I've seen a catalog layout that made sense for one team become the reason a second business unit couldn't get access cleanly. I've seen partitioning picked in week one that nobody looked at again until queries started timing out, and by then you're migrating data, not editing a config. Point the questions inward and they still land, because what they're really testing is the design.

A few of the 27 I'd weight higher than the article does.

The POC-versus-production one is where I've watched the most damage happen. A POC works, so it gets pushed to prod with no cluster policy, no auto-termination, no cost attribution by team. Everything looks fine until the first month-end DBU bill, and by then the design that generated the overrun already has pipelines sitting on top of it. When I ask this question I want to see the actual go-live checklist, not hear a definition of one.

Governance timing is the next one. The article calls "governance deferred" a disqualifier and I agree. Unity Catalog stood up after the data has already landed is one of the uglier things to unwind. Access, lineage, and classification either sit in the design from the start or you're booking a retrofit later.

The third is treating cost as an architecture property. The piece frames performance and cost as a trade-off, which is fair, but I'd put it more bluntly: on Databricks the bill is DBU consumption, so cost is mostly decided by architecture, not by watching a dashboard after the fact. Cluster policies, serverless versus job compute, storage maintenance, tagging for attribution. Those choices are the cost.

Two things I'd add.

The commercial questions near the end are written for someone hiring a firm, but they map onto an in-house build without much translation. Who owns the repo. Who holds workspace admin. Whether the effort estimate is honest about the assumptions under it. Same questions, pointed inward.

And one the article doesn't cover at all: how do you deploy changes after go-live? Asset Bundles or Terraform in a repo you control, or someone clicking through the UI on a Friday. That single answer decides whether your team can safely change anything six months in, and it almost never comes up during selection.

The part I'll keep is the distinction between a trade-off and a disqualifier. Wanting a local team or a particular pricing model is a preference. No named delivery team, governance pushed to phase two, no handover artifacts, no admin access to your own workspace are not preferences, and no polished demo should cover for them.

Databricks-image.png

My bet is that the deploy-after-go-live question is the highest-leverage one on the whole list and the one that gets skipped most in selection. Which do you see skipped most often?

0 REPLIES 0