cancel
Showing results forย 
Search instead forย 
Did you mean:ย 
Community Articles
Dive into a collaborative space where members like YOU can exchange knowledge, tips, and best practices. Join the conversation today and unlock a wealth of collective wisdom to enhance your experience and drive success.
cancel
Showing results forย 
Search instead forย 
Did you mean:ย 

Nexa: Genie is a Minute Away

ravikr1
Databricks Partner

Building an Explainable Semantic Intelligence Platform on the Databricks Lakehouse

We are building Nexa, a semantic intelligence platform for Databricks that turns a raw Lakehouse into a continuously maintained, explainable system that can compile natural language requests into governed Genie Agents. We are about 60 percent through implementation and want to share the architecture, the reasoning behind it, and why we think it solves a problem most catalog and semantic layer tools have not addressed properly: trust.

ravikr1_0-1789061751173.png

The problem we set out to solve

Every enterprise Lakehouse eventually accumulates the same gap. Unity Catalog tells you what tables and columns exist. Business users describe their world in a completely different vocabulary: revenue, shrink, available inventory, active customers. Someone, usually a data engineer, sits in the middle translating between the two, by hand, every single time a new question comes in.

Genie Agents solve part of this by letting business users ask questions in plain language. But a Genie Agent is only as good as the semantic mapping behind it, and most tools that generate that mapping treat the language model as the source of truth. That is a mistake. An LLM can guess that a column named inventory_qty probably means Available Inventory, but a guess is not a governed enterprise definition, and a wrong guess embedded silently in a Genie Agent produces confidently wrong answers that nobody catches until a business decision goes sideways.

Nexa is built on a different premise. AI can interpret technical truth, but it must never become the source of technical truth. Every semantic interpretation has to be traceable back to technical evidence, and every important AI decision has to be able to answer the question why do you believe this.

Two layers, kept honest

Nexa is architected as two distinct layers that are never allowed to merge into one blurry system.

ravikr1_0-1789062595929.png

 

The first is the Technical Knowledge Graph. This is the canonical, machine derived representation of the actual Databricks environment: schemas, tables, columns, primary and foreign key relationships, lineage, query behavior, data quality signals, tags, and physical statistics. Nothing in this layer is invented. It comes directly from Unity Catalog metadata, information schema, system tables, and observed lineage.

The second is the Enterprise Semantic Layer. This is the governed interpretation on top: business concepts, metrics, KPIs, dimensions, synonyms, and business rules. This layer is derived, not authoritative on its own. Every mapping from a business concept to a technical asset carries evidence, and that evidence is inspectable.

We deliberately separate two kinds of explainability that most platforms conflate. A Graph ML model answers why we believe two technical entities are related. A language model answers why we interpret a technical entity as a particular business concept. Keeping these two questions and their answers separate means that when something goes wrong, you know immediately whether the problem is structural or semantic, instead of debugging a black box.

ravikr1_1-1789062040145.png

 

How the pieces fit together

The system is organized into five planes.

The experience plane is a React application covering the Ontology Explorer, Semantic Explorer, graph visualization, Genie Builder, Genie Chat, governance review, explainability views, impact analysis, and drift and quality monitoring.

The application plane is a Node.js orchestrator handling authentication, the Databricks API client, the Genie Agent API client, job orchestration, semantic compiler orchestration, session management, and streaming.

Below that sit two engines working in parallel. The Python AI Engine handles profiling, embeddings, GraphSAGE and GAT based edge scoring, entity resolution, semantic matching, model evaluation, and MLflow tracking. The Semantic Compiler handles intent interpretation, concept resolution, metric selection, join planning, Genie configuration, instruction generation, benchmark generation, and quality gates.

Both feed into the Semantic Intelligence Plane, which holds the Technical Knowledge Graph, the Enterprise Semantic Layer, an Evidence Engine, a Trust Engine, an Ontology Critic, and the governance and human feedback loop.

At the bottom is the Data Truth Plane: Unity Catalog for information schema, table and column metadata, lineage, and tags, Delta Lake for ontology state, profiles, evidence, feedback, semantic definitions, and evaluation results, and Databricks system tables for operational signal.

We favor native Databricks capabilities wherever they provide reliable functionality. Unity Catalog metadata, lineage, and system tables get you most of the way to a technical graph before you need to invoke a single model. Graph ML gets reserved specifically for ambiguous relationship discovery, not as a default step, because deterministic signals are cheaper, faster, and more trustworthy than probabilistic ones whenever they are available.

What makes disagreement a feature, not a bug

One design decision we expect to be controversial in a good way: when Graph ML, the language model, lineage, and business rules disagree with each other, Nexa surfaces the conflict instead of picking a winner silently.

Here is a real pattern we designed around. Say a column inventory_qty gets mapped to a business concept. The graph confidence is 97 percent that this is a valid technical relationship. The LLM's semantic confidence is 84 percent that the mapping name is correct. But the same field has historically been used as both Available Inventory and On Hand Inventory in existing definitions elsewhere in the enterprise. Nexa detects this and marks it as requiring human review rather than auto approving it. We consider that outcome a success, not a failure. A platform that hides this kind of ambiguity behind a high confidence score is more dangerous than one that has no opinion at all.

This is enforced through what we call the fail closed principle. Low confidence or conflicting semantic decisions require human review rather than automatically becoming trusted enterprise definitions. High risk changes never sail through on autopilot.

Semantic Diff and Drift Timeline

Governance logs traditionally tell you that a decision happened. We wanted something better: a browsable history of what a concept has meant over time.

Every mapping between a business concept and a technical asset is stored as an append only version in a Delta table, never overwritten in place. Each version carries its graph confidence, its LLM confidence, its evidence reference, who approved it, and why. When two approved mappings for the same concept conflict, that gets logged separately with its own resolution workflow.

In the Semantic Explorer, this becomes a Concept Timeline: a horizontal sequence of every version a business concept has ever had, color coded by status, with a compare mode that shows exactly what changed between two versions and which Genie Agents would be affected by that change. If someone asks why the definition of Available Inventory changed last quarter, the answer is a few clicks away instead of a Slack archaeology project.

Counterfactual Genie: simulating impact before you commit

This is the feature we are most excited about, because we have not seen it anywhere else in the catalog or semantic layer space.

Before a semantic mapping, a schema change, or a confidence threshold adjustment gets approved, Nexa can simulate its effect through the graph and show exactly which live Genie Agents, dashboards, and metrics would change, and by how much, without ever touching the real system.

The mechanism is a shadow overlay rather than a shadow environment. We do not stand up a second graph or a second Genie deployment. Instead, a small hypothetical diff object gets passed as an optional parameter into the same graph read functions and the same semantic compiler code path that runs in production. When the overlay is present, reads get patched in memory before being returned. No write ever touches the real graph or the real Delta tables during a simulation. This guarantees the simulation can never drift from real behavior, because it is running the exact same logic, just against a hypothetical state.

In practice this means a reviewer looking at a proposed mapping change can click Simulate Impact and see, within the same review screen, a checklist of every Genie Agent that references the affected concept, whether each one would materially change, and a side by side diff of the generated query configuration before and after. If lowering an auto trust threshold would cause 23 additional mappings to auto approve last quarter, and 3 of them would contradict each other, that shows up before the threshold change is committed, not after.

We think of this as extending our fail closed principle into fail closed with a preview. Governance is not just a gate, it is a gate with a window.

Why this matters for enterprise

A business user should eventually be able to say build me a supply chain intelligence agent and have the platform work out what they mean, which business concepts and metrics are involved, which technical assets support those concepts, which relationships are trustworthy enough to use, what data should be exposed, how the Genie Agent should be configured, whether the result passes quality checks, and why every one of those decisions was made.

That last part, the why, is the piece most platforms skip. Nexa treats explainability and governance as core infrastructure rather than an afterthought layered on top of a text to SQL engine. The technical graph tells you what exists. The semantic layer tells you what it means. The evidence engine tells you why you should believe it. The governance layer tells you whether it is trusted. The semantic compiler turns all of that into a working Genie Agent, and the counterfactual simulator lets you see the consequences before you sign off on anything.

ravikr1_2-1789062130694.png

 

Where we are

We are roughly 60 percent through building Nexa end to end: the Technical Knowledge Graph ingestion from Unity Catalog, the Evidence Engine, and the core Semantic Compiler path are the furthest along, with the Semantic Diff timeline and Counterfactual Genie simulator actively being built out next.

Genie is a minute away. We built Nexa to make sure that minute is trustworthy, not just fast.

If you are working on similar problems around semantic layers, Genie Agent governance, or explainable AI on the Lakehouse, we would love to compare notes.

1 REPLY 1

ravikr1
Databricks Partner

Look at Visuals