Monday
As more teams start connecting AI agents and AI/BI workloads to Databricks, Iโve been thinking about what actually makes an enterprise data platform โAI-ready.โ
From my experience working in enterprise data and analytics at Polestar Analytics, the challenge often isn't just getting data into Delta tables. Itโs making sure the data is:
Well governed and access-controlled
Consistently defined across teams
Supported by reliable metadata and lineage
Structured around trusted business metrics
Easy for downstream AI systems and agents to understand
For example, an AI agent can generate technically valid SQL, but if โcustomer,โ โrevenue,โ or โactive accountโ means different things across datasets, the answer can still be wrong.
I'm curious how other Databricks teams are approaching this.
What do you consider the most important foundation for making a Databricks environment AI-ready โ data quality, Unity Catalog governance, metadata/lineage, semantic definitions, or something else?
Would be particularly interested in experiences from teams that have already moved beyond experimentation into production AI/agent use cases.
Tuesday
Hello, I think trusted and consistent business definitions are just as important as data quality and governance.
Wednesday
Agreed. I think the business-definition piece can easily get overlooked because data quality and governance are more obvious technical requirements.
A dataset can be clean and well governed, but if different teams use different definitions for something like โrevenueโ or โcustomer,โ the AI can still produce a technically correct but business-wrong answer.
Tuesday
Hello Karthik,
I would say AI-readiness starts with trusted and well-governed data rather than just focusing on the AI layer itself. In my view, data quality, Unity Catalog governance, data lineage, and consistent business definitions all need to work together and are interdepedent.
For AI/agent use cases especially, having clear semantic definitions and trusted business metrics is important, since technically valid queries can still produce incorrect answers when the underlying definitions are inconsistent.
Once these foundations and the requirements are in place, it becomes much easier to build reliable AI/BI and agent-based solutions on top of the platform.
Wednesday
Exactly. I like the point that these foundations are interdependent rather than separate checklist items.
The semantic-definition piece is particularly interesting for agent use cases. As agents move beyond answering questions and start taking actions, having consistent metrics, lineage and governance becomes even more important.
Curious if you've seen teams formalize those business definitions in a semantic layer or metric framework yet?
Wednesday
I agree with Aravind. Creating a data platform ready for AI analytics forces teams to ensure the basics are in place: documentation, data quality, certification, monitoring, etc. A lot of these are overlooked in hindsight but AI doesn't answer accurately without all of that at its disposal (unless you're massively overfitting). The foundations are key.
Wednesday
Hi Kartik, to your follow-up about whether teams have formalized business definitions in a semantic layer: on Databricks the closest native piece is Unity Catalog metric views. You define dimensions and measures once in YAML (source, joins, filter, dimensions, measures), register it in Unity Catalog, and query it with the MEASURE() function. The point for agents is that AI/BI dashboards, Genie and plain SQL can all resolve "revenue" or "active account" from the same governed definition, so the number cannot drift between the dashboard and the chat answer.
A few things I would add to the "what does AI-ready mean" list, because they are the ones that turn a checklist into something testable:
1) Treat definitions as code with an owner. Put the metric view YAML in Git, review changes like any other PR, and name one owner per metric. A common cause of wrong agent answers is two valid definitions of the same word, not bad SQL.
2) Build a small golden question set before you open it to users: 30 to 50 real questions with a reviewed SQL answer each. Run it after every change to a metric view, a table or the agent instructions. This is the closest thing to a regression test for "AI-ready", and it tells you quickly whether a schema change silently broke an agent.
3) Pick the grain deliberately. Agents are good at SELECT but poor at knowing which table is the trusted one when three look alike. Fewer, well-commented, certified tables with column descriptions help more than a large catalog with everything exposed.
4) Measure cost and answer quality per successful task, not only per query. Governance tells you who can touch what, but it does not tell you whether the agent got the right answer or how much it spent getting there.
On the metric-view side, one limit to know: measures must be queried with MEASURE() and SELECT * is not supported, so existing queries and tools that expect plain tables will need a small adaptation.
Curious whether the teams you have talked to start with the metrics or with the golden question set, because I would guess the second one exposes the first one's gaps fastest.
Wednesday
I see AI-readiness as more than making data accessible to an AI/agent. The foundation is trusted data + governed access + business context.
In Databricks, I would build this around Unity Catalog for governance, lineage and discoverability, strong data-quality controls in the Lakehouse, and a well-defined semantic/business layer so terms such as customer, revenue or active account have consistent meaning.
For production AI/agent workloads, I would also add observability and evaluation. It is not enough for an agent to generate valid SQL โ we need to know whether it used the right governed data, interpreted the business context correctly, and produced a trustworthy answer.
So for me: Governance + Data Quality + Semantic Context + Lineage + Observability/Evaluation = AI-ready data platform.