<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic The Missing Correlation Layer in Databricks ModelOps | Building an AI Health Control Plane for DBx in Community Articles</title>
    <link>https://community.databricks.com/t5/community-articles/the-missing-correlation-layer-in-databricks-modelops-building-an/m-p/168476#M1556</link>
    <description>&lt;H2&gt;Summary&lt;/H2&gt;&lt;P&gt;This post walks through the architecture of &lt;STRONG&gt;The Third Eye&lt;/STRONG&gt;, a continuous AI health and governed ModelOps control plane built entirely on native Databricks capabilities. The core constraint: &lt;STRONG&gt;it never recomputes a metric Databricks already computes.&lt;/STRONG&gt; It reads Lakehouse Monitoring's own output tables, the lineage system tables, Unity Gateway's usage and guardrail tables, and MLflow/UC registry objects, correlates across them, scores composite risk with business context, and drives a governed action loop on top. If you're running enough models that per-model dashboards have stopped being useful, this is the layer I think is missing from most Databricks AI/ML deployments.&lt;/P&gt;&lt;H2&gt;Scope boundary: decide this before writing any code&lt;/H2&gt;&lt;P&gt;The single most useful artifact in the design phase was writing down, explicitly, what Databricks already computes versus what actually needs building. Skipping this step is how governance projects balloon into duplicate monitoring stacks.&lt;/P&gt;&lt;DIV&gt;Signal Databricks-native source What it gives you What Third Eye does &lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD&gt;Drift / distribution stats&lt;/TD&gt;&lt;TD&gt;Lakehouse Monitoring (or Data Profiling): {output_schema}.{table}_profile_metrics, {output_schema}.{table}_drift_metrics&lt;/TD&gt;&lt;TD&gt;Per-column stats, consecutive and baseline drift, on a configured schedule&lt;/TD&gt;&lt;TD&gt;Read the tables directly. No PSI/KS reimplementation.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Model quality / accuracy&lt;/TD&gt;&lt;TD&gt;Lakehouse Monitoring, InferenceLog analysis type&lt;/TD&gt;&lt;TD&gt;Accuracy per model_id/version once ground truth is joined&lt;/TD&gt;&lt;TD&gt;Read from the profile table.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Lineage (table + column)&lt;/TD&gt;&lt;TD&gt;system.access.table_lineage, system.access.column_lineage, Lineage REST API&lt;/TD&gt;&lt;TD&gt;Automatic lineage across jobs, notebooks, pipelines, dashboards, DBSQL&lt;/TD&gt;&lt;TD&gt;Read directly for correlation. 1-year rolling retention on system tables (indefinite via Catalog Explorer/API since Sept 1, 2024); REST API returns one hop per call, so walk recursively for multi-hop.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Inference request/response logs&lt;/TD&gt;&lt;TD&gt;Classic inference tables, or the newer Unified Trace Table (Unity Gateway, OpenTelemetry, Beta)&lt;/TD&gt;&lt;TD&gt;Full request/response payloads, latency, status, model version served&lt;/TD&gt;&lt;TD&gt;Read. Prefer the Unified Trace Table for anything Gateway-routed.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Cost / usage / tokens&lt;/TD&gt;&lt;TD&gt;system.serving.served_entities, system.serving.endpoint_usage&lt;/TD&gt;&lt;TD&gt;Per-endpoint and per-served-entity usage, plus a usage_context map for custom attribution&lt;/TD&gt;&lt;TD&gt;Read directly for cost panels. No custom cost math.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;PII/PHI, unsafe content&lt;/TD&gt;&lt;TD&gt;Unity Gateway AI Guardrails&lt;/TD&gt;&lt;TD&gt;Detection/blocking/filtering at the gateway&lt;/TD&gt;&lt;TD&gt;Read violation events as a risk signal.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Model registry&lt;/TD&gt;&lt;TD&gt;MLflow Model Registry + UC model objects&lt;/TD&gt;&lt;TD&gt;Registered models, versions, aliases, experiments&lt;/TD&gt;&lt;TD&gt;Read via MLflow API / UC objects for the asset inventory.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;AI asset discovery&lt;/TD&gt;&lt;TD&gt;Unity Gateway AI Asset Registry&lt;/TD&gt;&lt;TD&gt;Catalog of governed models, agents, MCP servers, tools&lt;/TD&gt;&lt;TD&gt;Read as the primary zero-touch discovery source.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Fairness / bias&lt;/TD&gt;&lt;TD&gt;Lakehouse Monitoring fairness/bias support for classification models&lt;/TD&gt;&lt;TD&gt;Bias metrics on schedule, if configured&lt;/TD&gt;&lt;TD&gt;Read if configured; provision the monitor via API if not.&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;/DIV&gt;&lt;P&gt;&lt;STRONG&gt;The actual new engineering is seven things:&lt;/STRONG&gt; cross-signal correlation, criticality-weighted composite scoring, confidence scoring on top of that, LLM-generated root-cause narrative, zero-touch discovery/registration glue, a governed action layer (retrain, champion/challenger, gated promotion), and one unified multi-channel alert digest.&lt;/P&gt;&lt;H2&gt;Unity Catalog governance schema&lt;/H2&gt;&lt;P&gt;Everything lives in governance.model_health as Delta tables. The design principle: Third Eye's tables are pointers and derived state, not copies of Databricks' own data. The native tables stay the system of record.&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;CREATE CATALOG IF NOT EXISTS governance;
CREATE SCHEMA IF NOT EXISTS governance.model_health;

CREATE TABLE governance.model_health.model_registry_map (
  model_id STRING NOT NULL,
  model_name STRING,
  model_version STRING,
  serving_endpoint STRING,
  gateway_registered BOOLEAN,
  owning_team STRING,
  business_domain STRING,
  criticality_tier STRING,               -- config, defaulted from tag/domain
  lakehouse_monitor_configured BOOLEAN,
  created_at TIMESTAMP,
  is_active BOOLEAN
) USING DELTA;

CREATE TABLE governance.model_health.signal_index (
  model_id STRING,
  signal_type STRING,                    -- drift | accuracy | cost | usage | guardrail | lineage_change
  source_table STRING,                   -- fully qualified native table this reads from
  computed_at TIMESTAMP,
  latest_value DOUBLE,                   -- normalized numeric snapshot for scoring
  raw_reference STRING                   -- pointer back to the full native record
) USING DELTA
PARTITIONED BY (signal_type);

CREATE TABLE governance.model_health.lineage_events (
  model_id STRING,
  upstream_table STRING,
  event_type STRING,                     -- schema_change | new_write | etc
  event_time TIMESTAMP,&lt;/SPAN&gt;&lt;SPAN&gt;  entity_type STRING                     -- JOB | NOTEBOOK | PIPELINE | DASHBOARD_V3 | DBSQL_QUERY
) USING DELTA;

CREATE TABLE governance.model_health.risk_scores (
  model_id STRING,
  computed_at TIMESTAMP,
  drift_component DOUBLE,
  quality_component DOUBLE,
  cost_component DOUBLE,
  guardrail_component DOUBLE,
  criticality_weight DOUBLE,
  health_score DOUBLE,                   -- 0-100, weighted composite
  confidence DOUBLE,                     -- 0-1
  health_tier STRING                     -- healthy | watch | at_risk | critical
) USING DELTA;

CREATE TABLE governance.model_health.incidents (
  incident_id STRING NOT NULL,
  model_id STRING,
  opened_at TIMESTAMP,
  trigger_signals STRING,                -- JSON array of co-occurring signal_index rows
  lineage_context STRING,                -- JSON, linked lineage_events if in-window
  root_cause_narrative STRING,           -- LLM-generated
  root_cause_confidence DOUBLE,
  recommended_action STRING,             -- investigate | no_action | remediate
  status STRING                          -- open | acknowledged | resolved
) USING DELTA;

CREATE TABLE governance.model_health.remediation_suggestions (
  incident_id STRING,
  generated_at TIMESTAMP,
  explanation STRING,                    -- plain-language, from foundation model&lt;/SPAN&gt;&lt;SPAN&gt;  suggested_actions STRING,              -- JSON array, ranked by effort
  urgency STRING                         -- urgent | can_wait
) USING DELTA;

CREATE TABLE governance.model_health.risk_weights_config (
  criticality_tier STRING NOT NULL,
  w_drift DOUBLE,
  w_quality DOUBLE,
  w_cost DOUBLE,
  w_guardrail DOUBLE,
  criticality_weight DOUBLE
) USING DELTA;&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;H2&gt;Discovery/sync job (PySpark)&lt;/H2&gt;&lt;P&gt;Registering a model normally, via mlflow.register_model() or a UC model registration, should be the only onboarding step. An hourly job does the rest:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;python&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;from databricks.sdk import WorkspaceClient

def sync_model_registry(mlflow_client, gateway_client, w: WorkspaceClient):
    mlflow_models = mlflow_client.search_registered_models()
    gateway_assets = gateway_client.list_ai_asset_registry()
    known = spark.table("governance.model_health.model_registry_map") \
                 .select("model_id").collect()
    known_ids = {r.model_id for r in known}

    new_rows = []
    for m in mlflow_models:
        if m.model_id not in known_ids:
            new_rows.append(build_registry_row(m, gateway_assets))

    if new_rows:
        spark.createDataFrame(new_rows).write.mode("append") \
             .saveAsTable("governance.model_health.model_registry_map")

    for row in new_rows:
        lineage = walk_lineage(row["serving_endpoint"], hops=1)  # system table or REST API
        write_lineage_events(row["model_id"], lineage)
        if not row["lakehouse_monitor_configured"]:
            w.lakehouse_monitors.create(
                table_name=row["serving_endpoint"],
                assets_dir=f"/monitors/{row['model_id']}",
                output_schema_name="governance.model_health",
                inference_log=InferenceLogProfileType(
                    problem_type="regression",  # or classification
                    prediction_col="prediction",
                    timestamp_col="ts",
                    granularities=["1 day"],
                    model_id_col="model_version",&lt;/SPAN&gt;&lt;SPAN&gt;                ),
            )&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;The new model_registry_map row is the only thing that activates dashboard inclusion, Genie scope, and monitoring. The dashboard and Genie's metric views are parameterized off this table's contents, not hardcoded per model.&lt;/P&gt;&lt;H2&gt;Signal adapters&lt;/H2&gt;&lt;P&gt;Read-only normalizers that write into signal_index. Example for the drift adapter:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;python&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;def adapt_drift_signals(output_schema: str):
    drift_df = spark.table(f"{output_schema}.drift_metrics") \
        .filter(F.col("drift_type") == "consecutive") \
        .select(
            F.col("model_id"),
            F.lit("drift").alias("signal_type"),
            F.lit(f"{output_schema}.drift_metrics").alias("source_table"),
            F.current_timestamp().alias("computed_at"),
            F.col("js_distance").alias("latest_value"),   # or your chosen distance metric
            F.col("column_name").alias("raw_reference"),
        )
    drift_df.write.mode("append").saveAsTable("governance.model_health.signal_index")&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;The cost/usage and guardrail adapters follow the same shape, reading from system.serving.endpoint_usage and the Gateway guardrail event tables respectively. Normalize into the same six-column signal_index shape so the correlation engine below doesn't need to know which native table a signal came from.&lt;/P&gt;&lt;H2&gt;Correlation engine&lt;/H2&gt;&lt;P&gt;Runs as a Databricks Workflow task after each Lakehouse Monitoring refresh cycle. It's a time-window join, not a new statistical method:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;WITH recent_signals AS (
  SELECT *
  FROM governance.model_health.signal_index
  WHERE computed_at &amp;gt;= current_timestamp() - INTERVAL 2 HOURS
    AND latest_value &amp;gt; (
      SELECT threshold FROM governance.model_health.signal_thresholds t
      WHERE t.signal_type = signal_index.signal_type
    )
),
lineage_in_window AS (
  SELECT le.model_id, le.upstream_table, le.event_type, le.event_time
  FROM governance.model_health.lineage_events le
  JOIN recent_signals rs
    ON le.model_id = rs.model_id
   AND le.event_time BETWEEN rs.computed_at - INTERVAL 2 HOURS AND rs.computed_at
)
SELECT
  rs.model_id,
  to_json(collect_list(struct(rs.signal_type, rs.latest_value, rs.computed_at))) AS trigger_signals,
  to_json(collect_list(struct(lw.upstream_table, lw.event_type, lw.event_time))) AS lineage_context
FROM recent_signals rs
LEFT JOIN lineage_in_window lw ON rs.model_id = lw.model_id
GROUP BY rs.model_id
HAVING size(collect_list(lw.upstream_table)) &amp;gt; 0      -- signal + corroborating lineage event
    OR count(DISTINCT rs.signal_type) &amp;gt;= 2             -- or 2+ independent signal types agree&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;If a model has one weak signal and nothing corroborating it, this query produces no row for it, and no incident opens. That HAVING clause is the alert-fatigue control. It's as load-bearing as the join itself.&lt;/P&gt;&lt;H2&gt;Scoring engine (PySpark)&lt;/H2&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;python&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;from pyspark.sql import functions as F

def compute_health_score(signals_df, weights_df):
    joined = signals_df.join(weights_df, on="criticality_tier")
    return joined \
        .withColumn(
            "health_score",
            100 - (
                F.col("w_drift") * F.col("drift_component")
                + F.col("w_quality") * (1 - F.col("quality_component"))
                + F.col("w_cost") * F.col("cost_component")
                + F.col("w_guardrail") * F.col("guardrail_component")
            ) * F.col("criticality_weight")
        ) \
        .withColumn(
            "confidence",
            F.least(
                F.col("evidence_volume") / F.lit(30.0),
                F.col("baseline_window_completeness"),
                F.when(F.col("ground_truth_available"), F.lit(1.0)).otherwise(F.lit(0.6)),
            )
        ) \
        .withColumn(
            "health_tier",
            F.when(F.col("health_score") &amp;gt;= 85, "healthy")
             .when(F.col("health_score") &amp;gt;= 65, "watch")
             .when(F.col("health_score") &amp;gt;= 40, "at_risk")
             .otherwise("critical")
        )&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;w_drift, w_quality, w_cost, w_guardrail, and criticality_weight all come from risk_weights_config, keyed on tier. Never hardcoded, so a large drift on a Tier 3 experimental model can rank below a small drift on a Tier 1 regulated one.&lt;/P&gt;&lt;H2&gt;Root-cause narrative via ai_query()&lt;/H2&gt;&lt;P&gt;Once an incident opens, a Workflow task calls a Foundation Model endpoint with the evidence bundle, entirely inside Databricks:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;SELECT
  incident_id,
  ai_query(
    'databricks-meta-llama-3-70b-instruct',
    concat(
      'Evidence bundle: ', trigger_signals, ' Lineage context: ', lineage_context, '. ',
      'In 3 sentences: explain the likely root cause, ',
      'list remediation actions ranked by effort, ',
      'and classify urgency as urgent or can_wait.'
    )
  ) AS explanation
FROM governance.model_health.incidents
WHERE status = 'open'
  AND incident_id NOT IN (SELECT incident_id FROM governance.model_health.remediation_suggestions)&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;Write the result to remediation_suggestions. The dashboard, Genie, and Copilot Studio are pure readers of this table. None of them re-derive or reformat the explanation independently, which avoids three slightly different stories about the same incident.&lt;/P&gt;&lt;H2&gt;Serving layer&lt;/H2&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;AI/BI Dashboard&lt;/STRONG&gt;: Fleet Overview (KPI tiles for % healthy/watch/at_risk/critical, health_score trend, open incident count), Model Leaderboard (sortable by health_score/criticality, filterable by domain/team, drift sparkline, last-retrain date, cost trend), and a Model Detail drill-through (health_score time series with component breakdown, feature-level drift pulled straight from the native Lakehouse Monitoring table, a lineage snippet, open/past incidents with root-cause narrative, cost/usage trend).&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Genie Space&lt;/STRONG&gt;: metric views over risk_scores, incidents, and signal_index, plus an instruction doc mapping business vocabulary ("risky," "drifting," "stale," "costly") to specific columns/thresholds so Genie resolves questions without ad hoc joins per query. A metric view definition looks like:&lt;/LI&gt;&lt;/UL&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;CREATE VIEW governance.model_health.vw_fleet_risk AS
SELECT
  m.model_name, m.business_domain, m.criticality_tier,
  r.health_score, r.health_tier, r.confidence, r.computed_at
FROM governance.model_health.risk_scores r
JOIN governance.model_health.model_registry_map m USING (model_id)
QUALIFY ROW_NUMBER() OVER (PARTITION BY m.model_id ORDER BY r.computed_at DESC) = 1;&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Copilot Studio&lt;/STRONG&gt;: two message types, both pure readers of the governed schema. A scheduled KPI digest posted to Teams and fanned out to Slack/Google Chat via webhook actions in the same topic, and a breach notification triggered by an incidents creation event, reading the linked remediation_suggestions row and formatting it as an adaptive card.&lt;/LI&gt;&lt;/UL&gt;&lt;H2&gt;Platform limits to design around&lt;/H2&gt;&lt;UL&gt;&lt;LI&gt;Lineage isn't preserved across renames of catalogs/schemas/tables/columns, and doesn't exist before September 1, 2024.&lt;/LI&gt;&lt;LI&gt;Lineage system tables carry a rolling 1-year retention (Catalog Explorer/API retain indefinitely since Sept 2024), so snapshot anything needed for long-term trend charts into your own tables.&lt;/LI&gt;&lt;LI&gt;The Lineage REST API returns one hop per call. Walk recursively for multi-hop graphs, or query the system tables directly for a full join.&lt;/LI&gt;&lt;LI&gt;Trace/inference table delivery for Gateway-routed traffic is best-effort, with delays up to roughly an hour, and inference tables aren't guaranteed for 401/403/429/500 responses.&lt;/LI&gt;&lt;/UL&gt;&lt;H2&gt;Build order&lt;/H2&gt;&lt;OL&gt;&lt;LI&gt;Scaffolding: the DDL above, plus risk_weights_config and criticality defaults.&lt;/LI&gt;&lt;LI&gt;Discovery/sync job: build against a synthetic/mock registry first if you don't have live Gateway access yet.&lt;/LI&gt;&lt;LI&gt;Signal adapters: the largest glue surface, so budget the most time here.&lt;/LI&gt;&lt;LI&gt;Correlation and scoring engine: the one truly novel computation, and it deserves the most test coverage.&lt;/LI&gt;&lt;LI&gt;Root-cause narrative job.&lt;/LI&gt;&lt;LI&gt;AI/BI dashboard (three pages), Genie Space (metric views plus instruction doc), alerting plus Copilot Studio topics.&lt;/LI&gt;&lt;LI&gt;Model passport generator: a formatted read of the governed schema per model. Cheap, high demo value.&lt;/LI&gt;&lt;LI&gt;README, architecture diagram, and a demo script that injects a synthetic drift/lineage-change scenario end to end: discovery, then correlation, then dashboard update, then Genie investigation, then a Teams alert with remediation.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Auto-retraining, blast-radius analysis, and multi-workspace federation share the same data model and scoring engine but don't start until the above runs clean end to end on synthetic data.&lt;/P&gt;&lt;H2&gt;Tech stack&lt;/H2&gt;&lt;P&gt;Unity Catalog and Delta Lake; Lakehouse Monitoring / Data Profiling; system.access.table_lineage and column_lineage; Unity Gateway (AI Asset Registry, Unified Trace Table, system.serving.*, AI Guardrails); MLflow Model Registry; Databricks Workflows; Foundation Model APIs via ai_query(); Databricks AI/BI (Lakeview) dashboards; Genie Space with metric views; Microsoft Copilot Studio fanned out to Slack/Google Chat; PySpark and Databricks SQL for the adapters and scoring jobs.&lt;/P&gt;&lt;P&gt;I'd welcome feedback from anyone running Lakehouse Monitoring or Unity Gateway at fleet scale, particularly on the correlation window sizing in the query above, and on whether the signal_index pointer-table pattern holds up against your own retention requirements.&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 09_00_11 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31055i44C72AFDB4ABA3EF/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 09_00_11 PM.png" alt="ChatGPT Image Sep 13, 2026, 09_00_11 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_54_12 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31053iB67ABA6B0F8E6617/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_54_12 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_54_12 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_52_09 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31056i8C6E47C5A72755A0/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_52_09 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_52_09 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_49_11 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31054i923D45EABD0794BD/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_49_11 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_49_11 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_47_53 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31057i25E4108D56FCC4D3/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_47_53 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_47_53 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;</description>
    <pubDate>Sun, 13 Sep 2026 15:56:20 GMT</pubDate>
    <dc:creator>ravikr1</dc:creator>
    <dc:date>2026-09-13T15:56:20Z</dc:date>
    <item>
      <title>The Missing Correlation Layer in Databricks ModelOps | Building an AI Health Control Plane for DBx</title>
      <link>https://community.databricks.com/t5/community-articles/the-missing-correlation-layer-in-databricks-modelops-building-an/m-p/168476#M1556</link>
      <description>&lt;H2&gt;Summary&lt;/H2&gt;&lt;P&gt;This post walks through the architecture of &lt;STRONG&gt;The Third Eye&lt;/STRONG&gt;, a continuous AI health and governed ModelOps control plane built entirely on native Databricks capabilities. The core constraint: &lt;STRONG&gt;it never recomputes a metric Databricks already computes.&lt;/STRONG&gt; It reads Lakehouse Monitoring's own output tables, the lineage system tables, Unity Gateway's usage and guardrail tables, and MLflow/UC registry objects, correlates across them, scores composite risk with business context, and drives a governed action loop on top. If you're running enough models that per-model dashboards have stopped being useful, this is the layer I think is missing from most Databricks AI/ML deployments.&lt;/P&gt;&lt;H2&gt;Scope boundary: decide this before writing any code&lt;/H2&gt;&lt;P&gt;The single most useful artifact in the design phase was writing down, explicitly, what Databricks already computes versus what actually needs building. Skipping this step is how governance projects balloon into duplicate monitoring stacks.&lt;/P&gt;&lt;DIV&gt;Signal Databricks-native source What it gives you What Third Eye does &lt;TABLE&gt;&lt;TBODY&gt;&lt;TR&gt;&lt;TD&gt;Drift / distribution stats&lt;/TD&gt;&lt;TD&gt;Lakehouse Monitoring (or Data Profiling): {output_schema}.{table}_profile_metrics, {output_schema}.{table}_drift_metrics&lt;/TD&gt;&lt;TD&gt;Per-column stats, consecutive and baseline drift, on a configured schedule&lt;/TD&gt;&lt;TD&gt;Read the tables directly. No PSI/KS reimplementation.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Model quality / accuracy&lt;/TD&gt;&lt;TD&gt;Lakehouse Monitoring, InferenceLog analysis type&lt;/TD&gt;&lt;TD&gt;Accuracy per model_id/version once ground truth is joined&lt;/TD&gt;&lt;TD&gt;Read from the profile table.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Lineage (table + column)&lt;/TD&gt;&lt;TD&gt;system.access.table_lineage, system.access.column_lineage, Lineage REST API&lt;/TD&gt;&lt;TD&gt;Automatic lineage across jobs, notebooks, pipelines, dashboards, DBSQL&lt;/TD&gt;&lt;TD&gt;Read directly for correlation. 1-year rolling retention on system tables (indefinite via Catalog Explorer/API since Sept 1, 2024); REST API returns one hop per call, so walk recursively for multi-hop.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Inference request/response logs&lt;/TD&gt;&lt;TD&gt;Classic inference tables, or the newer Unified Trace Table (Unity Gateway, OpenTelemetry, Beta)&lt;/TD&gt;&lt;TD&gt;Full request/response payloads, latency, status, model version served&lt;/TD&gt;&lt;TD&gt;Read. Prefer the Unified Trace Table for anything Gateway-routed.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Cost / usage / tokens&lt;/TD&gt;&lt;TD&gt;system.serving.served_entities, system.serving.endpoint_usage&lt;/TD&gt;&lt;TD&gt;Per-endpoint and per-served-entity usage, plus a usage_context map for custom attribution&lt;/TD&gt;&lt;TD&gt;Read directly for cost panels. No custom cost math.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;PII/PHI, unsafe content&lt;/TD&gt;&lt;TD&gt;Unity Gateway AI Guardrails&lt;/TD&gt;&lt;TD&gt;Detection/blocking/filtering at the gateway&lt;/TD&gt;&lt;TD&gt;Read violation events as a risk signal.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Model registry&lt;/TD&gt;&lt;TD&gt;MLflow Model Registry + UC model objects&lt;/TD&gt;&lt;TD&gt;Registered models, versions, aliases, experiments&lt;/TD&gt;&lt;TD&gt;Read via MLflow API / UC objects for the asset inventory.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;AI asset discovery&lt;/TD&gt;&lt;TD&gt;Unity Gateway AI Asset Registry&lt;/TD&gt;&lt;TD&gt;Catalog of governed models, agents, MCP servers, tools&lt;/TD&gt;&lt;TD&gt;Read as the primary zero-touch discovery source.&lt;/TD&gt;&lt;/TR&gt;&lt;TR&gt;&lt;TD&gt;Fairness / bias&lt;/TD&gt;&lt;TD&gt;Lakehouse Monitoring fairness/bias support for classification models&lt;/TD&gt;&lt;TD&gt;Bias metrics on schedule, if configured&lt;/TD&gt;&lt;TD&gt;Read if configured; provision the monitor via API if not.&lt;/TD&gt;&lt;/TR&gt;&lt;/TBODY&gt;&lt;/TABLE&gt;&lt;/DIV&gt;&lt;P&gt;&lt;STRONG&gt;The actual new engineering is seven things:&lt;/STRONG&gt; cross-signal correlation, criticality-weighted composite scoring, confidence scoring on top of that, LLM-generated root-cause narrative, zero-touch discovery/registration glue, a governed action layer (retrain, champion/challenger, gated promotion), and one unified multi-channel alert digest.&lt;/P&gt;&lt;H2&gt;Unity Catalog governance schema&lt;/H2&gt;&lt;P&gt;Everything lives in governance.model_health as Delta tables. The design principle: Third Eye's tables are pointers and derived state, not copies of Databricks' own data. The native tables stay the system of record.&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;CREATE CATALOG IF NOT EXISTS governance;
CREATE SCHEMA IF NOT EXISTS governance.model_health;

CREATE TABLE governance.model_health.model_registry_map (
  model_id STRING NOT NULL,
  model_name STRING,
  model_version STRING,
  serving_endpoint STRING,
  gateway_registered BOOLEAN,
  owning_team STRING,
  business_domain STRING,
  criticality_tier STRING,               -- config, defaulted from tag/domain
  lakehouse_monitor_configured BOOLEAN,
  created_at TIMESTAMP,
  is_active BOOLEAN
) USING DELTA;

CREATE TABLE governance.model_health.signal_index (
  model_id STRING,
  signal_type STRING,                    -- drift | accuracy | cost | usage | guardrail | lineage_change
  source_table STRING,                   -- fully qualified native table this reads from
  computed_at TIMESTAMP,
  latest_value DOUBLE,                   -- normalized numeric snapshot for scoring
  raw_reference STRING                   -- pointer back to the full native record
) USING DELTA
PARTITIONED BY (signal_type);

CREATE TABLE governance.model_health.lineage_events (
  model_id STRING,
  upstream_table STRING,
  event_type STRING,                     -- schema_change | new_write | etc
  event_time TIMESTAMP,&lt;/SPAN&gt;&lt;SPAN&gt;  entity_type STRING                     -- JOB | NOTEBOOK | PIPELINE | DASHBOARD_V3 | DBSQL_QUERY
) USING DELTA;

CREATE TABLE governance.model_health.risk_scores (
  model_id STRING,
  computed_at TIMESTAMP,
  drift_component DOUBLE,
  quality_component DOUBLE,
  cost_component DOUBLE,
  guardrail_component DOUBLE,
  criticality_weight DOUBLE,
  health_score DOUBLE,                   -- 0-100, weighted composite
  confidence DOUBLE,                     -- 0-1
  health_tier STRING                     -- healthy | watch | at_risk | critical
) USING DELTA;

CREATE TABLE governance.model_health.incidents (
  incident_id STRING NOT NULL,
  model_id STRING,
  opened_at TIMESTAMP,
  trigger_signals STRING,                -- JSON array of co-occurring signal_index rows
  lineage_context STRING,                -- JSON, linked lineage_events if in-window
  root_cause_narrative STRING,           -- LLM-generated
  root_cause_confidence DOUBLE,
  recommended_action STRING,             -- investigate | no_action | remediate
  status STRING                          -- open | acknowledged | resolved
) USING DELTA;

CREATE TABLE governance.model_health.remediation_suggestions (
  incident_id STRING,
  generated_at TIMESTAMP,
  explanation STRING,                    -- plain-language, from foundation model&lt;/SPAN&gt;&lt;SPAN&gt;  suggested_actions STRING,              -- JSON array, ranked by effort
  urgency STRING                         -- urgent | can_wait
) USING DELTA;

CREATE TABLE governance.model_health.risk_weights_config (
  criticality_tier STRING NOT NULL,
  w_drift DOUBLE,
  w_quality DOUBLE,
  w_cost DOUBLE,
  w_guardrail DOUBLE,
  criticality_weight DOUBLE
) USING DELTA;&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;H2&gt;Discovery/sync job (PySpark)&lt;/H2&gt;&lt;P&gt;Registering a model normally, via mlflow.register_model() or a UC model registration, should be the only onboarding step. An hourly job does the rest:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;python&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;from databricks.sdk import WorkspaceClient

def sync_model_registry(mlflow_client, gateway_client, w: WorkspaceClient):
    mlflow_models = mlflow_client.search_registered_models()
    gateway_assets = gateway_client.list_ai_asset_registry()
    known = spark.table("governance.model_health.model_registry_map") \
                 .select("model_id").collect()
    known_ids = {r.model_id for r in known}

    new_rows = []
    for m in mlflow_models:
        if m.model_id not in known_ids:
            new_rows.append(build_registry_row(m, gateway_assets))

    if new_rows:
        spark.createDataFrame(new_rows).write.mode("append") \
             .saveAsTable("governance.model_health.model_registry_map")

    for row in new_rows:
        lineage = walk_lineage(row["serving_endpoint"], hops=1)  # system table or REST API
        write_lineage_events(row["model_id"], lineage)
        if not row["lakehouse_monitor_configured"]:
            w.lakehouse_monitors.create(
                table_name=row["serving_endpoint"],
                assets_dir=f"/monitors/{row['model_id']}",
                output_schema_name="governance.model_health",
                inference_log=InferenceLogProfileType(
                    problem_type="regression",  # or classification
                    prediction_col="prediction",
                    timestamp_col="ts",
                    granularities=["1 day"],
                    model_id_col="model_version",&lt;/SPAN&gt;&lt;SPAN&gt;                ),
            )&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;The new model_registry_map row is the only thing that activates dashboard inclusion, Genie scope, and monitoring. The dashboard and Genie's metric views are parameterized off this table's contents, not hardcoded per model.&lt;/P&gt;&lt;H2&gt;Signal adapters&lt;/H2&gt;&lt;P&gt;Read-only normalizers that write into signal_index. Example for the drift adapter:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;python&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;def adapt_drift_signals(output_schema: str):
    drift_df = spark.table(f"{output_schema}.drift_metrics") \
        .filter(F.col("drift_type") == "consecutive") \
        .select(
            F.col("model_id"),
            F.lit("drift").alias("signal_type"),
            F.lit(f"{output_schema}.drift_metrics").alias("source_table"),
            F.current_timestamp().alias("computed_at"),
            F.col("js_distance").alias("latest_value"),   # or your chosen distance metric
            F.col("column_name").alias("raw_reference"),
        )
    drift_df.write.mode("append").saveAsTable("governance.model_health.signal_index")&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;The cost/usage and guardrail adapters follow the same shape, reading from system.serving.endpoint_usage and the Gateway guardrail event tables respectively. Normalize into the same six-column signal_index shape so the correlation engine below doesn't need to know which native table a signal came from.&lt;/P&gt;&lt;H2&gt;Correlation engine&lt;/H2&gt;&lt;P&gt;Runs as a Databricks Workflow task after each Lakehouse Monitoring refresh cycle. It's a time-window join, not a new statistical method:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;WITH recent_signals AS (
  SELECT *
  FROM governance.model_health.signal_index
  WHERE computed_at &amp;gt;= current_timestamp() - INTERVAL 2 HOURS
    AND latest_value &amp;gt; (
      SELECT threshold FROM governance.model_health.signal_thresholds t
      WHERE t.signal_type = signal_index.signal_type
    )
),
lineage_in_window AS (
  SELECT le.model_id, le.upstream_table, le.event_type, le.event_time
  FROM governance.model_health.lineage_events le
  JOIN recent_signals rs
    ON le.model_id = rs.model_id
   AND le.event_time BETWEEN rs.computed_at - INTERVAL 2 HOURS AND rs.computed_at
)
SELECT
  rs.model_id,
  to_json(collect_list(struct(rs.signal_type, rs.latest_value, rs.computed_at))) AS trigger_signals,
  to_json(collect_list(struct(lw.upstream_table, lw.event_type, lw.event_time))) AS lineage_context
FROM recent_signals rs
LEFT JOIN lineage_in_window lw ON rs.model_id = lw.model_id
GROUP BY rs.model_id
HAVING size(collect_list(lw.upstream_table)) &amp;gt; 0      -- signal + corroborating lineage event
    OR count(DISTINCT rs.signal_type) &amp;gt;= 2             -- or 2+ independent signal types agree&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;If a model has one weak signal and nothing corroborating it, this query produces no row for it, and no incident opens. That HAVING clause is the alert-fatigue control. It's as load-bearing as the join itself.&lt;/P&gt;&lt;H2&gt;Scoring engine (PySpark)&lt;/H2&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;python&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;from pyspark.sql import functions as F

def compute_health_score(signals_df, weights_df):
    joined = signals_df.join(weights_df, on="criticality_tier")
    return joined \
        .withColumn(
            "health_score",
            100 - (
                F.col("w_drift") * F.col("drift_component")
                + F.col("w_quality") * (1 - F.col("quality_component"))
                + F.col("w_cost") * F.col("cost_component")
                + F.col("w_guardrail") * F.col("guardrail_component")
            ) * F.col("criticality_weight")
        ) \
        .withColumn(
            "confidence",
            F.least(
                F.col("evidence_volume") / F.lit(30.0),
                F.col("baseline_window_completeness"),
                F.when(F.col("ground_truth_available"), F.lit(1.0)).otherwise(F.lit(0.6)),
            )
        ) \
        .withColumn(
            "health_tier",
            F.when(F.col("health_score") &amp;gt;= 85, "healthy")
             .when(F.col("health_score") &amp;gt;= 65, "watch")
             .when(F.col("health_score") &amp;gt;= 40, "at_risk")
             .otherwise("critical")
        )&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;w_drift, w_quality, w_cost, w_guardrail, and criticality_weight all come from risk_weights_config, keyed on tier. Never hardcoded, so a large drift on a Tier 3 experimental model can rank below a small drift on a Tier 1 regulated one.&lt;/P&gt;&lt;H2&gt;Root-cause narrative via ai_query()&lt;/H2&gt;&lt;P&gt;Once an incident opens, a Workflow task calls a Foundation Model endpoint with the evidence bundle, entirely inside Databricks:&lt;/P&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;SELECT
  incident_id,
  ai_query(
    'databricks-meta-llama-3-70b-instruct',
    concat(
      'Evidence bundle: ', trigger_signals, ' Lineage context: ', lineage_context, '. ',
      'In 3 sentences: explain the likely root cause, ',
      'list remediation actions ranked by effort, ',
      'and classify urgency as urgent or can_wait.'
    )
  ) AS explanation
FROM governance.model_health.incidents
WHERE status = 'open'
  AND incident_id NOT IN (SELECT incident_id FROM governance.model_health.remediation_suggestions)&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;Write the result to remediation_suggestions. The dashboard, Genie, and Copilot Studio are pure readers of this table. None of them re-derive or reformat the explanation independently, which avoids three slightly different stories about the same incident.&lt;/P&gt;&lt;H2&gt;Serving layer&lt;/H2&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;AI/BI Dashboard&lt;/STRONG&gt;: Fleet Overview (KPI tiles for % healthy/watch/at_risk/critical, health_score trend, open incident count), Model Leaderboard (sortable by health_score/criticality, filterable by domain/team, drift sparkline, last-retrain date, cost trend), and a Model Detail drill-through (health_score time series with component breakdown, feature-level drift pulled straight from the native Lakehouse Monitoring table, a lineage snippet, open/past incidents with root-cause narrative, cost/usage trend).&lt;/LI&gt;&lt;LI&gt;&lt;STRONG&gt;Genie Space&lt;/STRONG&gt;: metric views over risk_scores, incidents, and signal_index, plus an instruction doc mapping business vocabulary ("risky," "drifting," "stale," "costly") to specific columns/thresholds so Genie resolves questions without ad hoc joins per query. A metric view definition looks like:&lt;/LI&gt;&lt;/UL&gt;&lt;DIV&gt;&lt;DIV&gt;&lt;DIV&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV&gt;sql&lt;/DIV&gt;&lt;DIV&gt;&lt;PRE&gt;&lt;SPAN&gt;CREATE VIEW governance.model_health.vw_fleet_risk AS
SELECT
  m.model_name, m.business_domain, m.criticality_tier,
  r.health_score, r.health_tier, r.confidence, r.computed_at
FROM governance.model_health.risk_scores r
JOIN governance.model_health.model_registry_map m USING (model_id)
QUALIFY ROW_NUMBER() OVER (PARTITION BY m.model_id ORDER BY r.computed_at DESC) = 1;&lt;/SPAN&gt;&lt;/PRE&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Copilot Studio&lt;/STRONG&gt;: two message types, both pure readers of the governed schema. A scheduled KPI digest posted to Teams and fanned out to Slack/Google Chat via webhook actions in the same topic, and a breach notification triggered by an incidents creation event, reading the linked remediation_suggestions row and formatting it as an adaptive card.&lt;/LI&gt;&lt;/UL&gt;&lt;H2&gt;Platform limits to design around&lt;/H2&gt;&lt;UL&gt;&lt;LI&gt;Lineage isn't preserved across renames of catalogs/schemas/tables/columns, and doesn't exist before September 1, 2024.&lt;/LI&gt;&lt;LI&gt;Lineage system tables carry a rolling 1-year retention (Catalog Explorer/API retain indefinitely since Sept 2024), so snapshot anything needed for long-term trend charts into your own tables.&lt;/LI&gt;&lt;LI&gt;The Lineage REST API returns one hop per call. Walk recursively for multi-hop graphs, or query the system tables directly for a full join.&lt;/LI&gt;&lt;LI&gt;Trace/inference table delivery for Gateway-routed traffic is best-effort, with delays up to roughly an hour, and inference tables aren't guaranteed for 401/403/429/500 responses.&lt;/LI&gt;&lt;/UL&gt;&lt;H2&gt;Build order&lt;/H2&gt;&lt;OL&gt;&lt;LI&gt;Scaffolding: the DDL above, plus risk_weights_config and criticality defaults.&lt;/LI&gt;&lt;LI&gt;Discovery/sync job: build against a synthetic/mock registry first if you don't have live Gateway access yet.&lt;/LI&gt;&lt;LI&gt;Signal adapters: the largest glue surface, so budget the most time here.&lt;/LI&gt;&lt;LI&gt;Correlation and scoring engine: the one truly novel computation, and it deserves the most test coverage.&lt;/LI&gt;&lt;LI&gt;Root-cause narrative job.&lt;/LI&gt;&lt;LI&gt;AI/BI dashboard (three pages), Genie Space (metric views plus instruction doc), alerting plus Copilot Studio topics.&lt;/LI&gt;&lt;LI&gt;Model passport generator: a formatted read of the governed schema per model. Cheap, high demo value.&lt;/LI&gt;&lt;LI&gt;README, architecture diagram, and a demo script that injects a synthetic drift/lineage-change scenario end to end: discovery, then correlation, then dashboard update, then Genie investigation, then a Teams alert with remediation.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Auto-retraining, blast-radius analysis, and multi-workspace federation share the same data model and scoring engine but don't start until the above runs clean end to end on synthetic data.&lt;/P&gt;&lt;H2&gt;Tech stack&lt;/H2&gt;&lt;P&gt;Unity Catalog and Delta Lake; Lakehouse Monitoring / Data Profiling; system.access.table_lineage and column_lineage; Unity Gateway (AI Asset Registry, Unified Trace Table, system.serving.*, AI Guardrails); MLflow Model Registry; Databricks Workflows; Foundation Model APIs via ai_query(); Databricks AI/BI (Lakeview) dashboards; Genie Space with metric views; Microsoft Copilot Studio fanned out to Slack/Google Chat; PySpark and Databricks SQL for the adapters and scoring jobs.&lt;/P&gt;&lt;P&gt;I'd welcome feedback from anyone running Lakehouse Monitoring or Unity Gateway at fleet scale, particularly on the correlation window sizing in the query above, and on whether the signal_index pointer-table pattern holds up against your own retention requirements.&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 09_00_11 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31055i44C72AFDB4ABA3EF/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 09_00_11 PM.png" alt="ChatGPT Image Sep 13, 2026, 09_00_11 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_54_12 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31053iB67ABA6B0F8E6617/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_54_12 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_54_12 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_52_09 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31056i8C6E47C5A72755A0/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_52_09 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_52_09 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_49_11 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31054i923D45EABD0794BD/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_49_11 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_49_11 PM.png" /&gt;&lt;/span&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="ChatGPT Image Sep 13, 2026, 08_47_53 PM.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/31057i25E4108D56FCC4D3/image-size/large?v=v2&amp;amp;px=999" role="button" title="ChatGPT Image Sep 13, 2026, 08_47_53 PM.png" alt="ChatGPT Image Sep 13, 2026, 08_47_53 PM.png" /&gt;&lt;/span&gt;&lt;/P&gt;</description>
      <pubDate>Sun, 13 Sep 2026 15:56:20 GMT</pubDate>
      <guid>https://community.databricks.com/t5/community-articles/the-missing-correlation-layer-in-databricks-modelops-building-an/m-p/168476#M1556</guid>
      <dc:creator>ravikr1</dc:creator>
      <dc:date>2026-09-13T15:56:20Z</dc:date>
    </item>
  </channel>
</rss>

