Hi,
Great question. This is where most Medallion implementations either scale or get messy. Here's the pattern that has worked well, based on Kimball conformed dimensions mapped onto Unity Catalog.
Quick note on naming: in Databricks, "Zerobus" is a product (Zerobus Ingest, a streaming ingestion API in Lakeflow Connect). Calling this the Kimball Bus Architecture avoids confusion in searches and discussions.
1. Creating the Golden Record
Don't add a separate layer before Silver. Do the conforming in Silver, as its own zone:
Bronze: raw data, one table set per source system (CRM, ERP, e-commerceโฆ).
Silver (per source): cleansed and typed, still one set per source.
Silver/Conformed: where the golden record gets built. This step does:
Identity resolution: a crosswalk (key map) table, (source_system, source_key) โ customer_sk.
Survivorship rules: which source wins for which attribute. Versioned in Git.
Surrogate keys: deterministic hash keys (e.g. xxhash64 or sha2 of the business key) are reproducible and avoid concurrency problems with identity columns.
SCD history: Lakeflow Declarative Pipelines' AUTO CDC API handles SCD Type 1/2 and out-of-order events for you.
Gold: domain marts that consume the conformed dimensions and never re-derive them.
The key rule is one writer per conformed dimension: one pipeline, one owning team, one table. Every Gold mart references it and never copies or rebuilds it.
If you already run a real MDM tool (Reltio, Informatica, Profisee, etc.), treat its output as a source. Don't rebuild MDM in Spark.
Also add unknown / inferred member rows (e.g. sk = -1) so facts that arrive before their dimension rows still load and don't get dropped.
2. Governance & Unity Catalog Structure
A layout that isolates shared dimensions from domain logic:
prod_silver.<source>.* -- per-source cleansed data
prod_conformed.dims.* -- physical conformed dims (platform team writes)
prod_conformed.published.* -- views = the public contract
prod_gold.sales.* -- domain marts (sales team)
prod_gold.finance.* -- domain marts (finance team)
A dedicated catalog or schema for conformed dimensions. Only the owning team can write to it. Domain teams get SELECT on the published views only.
Views as the contract layer. Consumers query published.dim_customer, a view over the physical table. You can then change physical storage without breaking anyone.
Central security. Put row filters and column masks on the conformed dimensions once, and every mart inherits the PII protection.
Catalogs per environment (dev/test/prod), with workspaceโcatalog binding so dev jobs can't touch prod.
Tags for owner, classification and deprecation status.
Lineage (system.access.table_lineage / column_lineage) to see exactly who consumes each dimension. You'll need this for change management.
Delta Sharing if marts live in another metastore or region.
3. Change Management & Evolution
Treat each conformed dimension as a data product with a contract: grain, keys, schema, SLA and owner, with the DDL in Git, deployed through CI/CD (Asset Bundles or Terraform).
Non-breaking changes (allowed anytime):
Adding nullable columns.
Widening types (int โ bigint, etc.) via delta.enableTypeWidening, without rewriting data.
Breaking changes (renames, drops, grain or key changes):
Build dim_customer_v2 next to v1.
Keep published.dim_customer pointing to v1, and expose published.dim_customer_v2.
Use lineage to find all downstream consumers and notify their owners.
Tag v1 as deprecated with a sunset date. Migrate consumers, then switch the view and drop v1.
Column mapping (delta.columnMapping.mode = 'name') lets you rename or drop physically without a rewrite. For consumers, though, a rename is still a breaking change, so do it behind the view.
Guardrails that prevent incidents:
WriteโAuditโPublish: build into a staging table and run expectations (unique SK, non-null business key, exactly one current row per key in SCD2). Publish only if they pass.
No SELECT * and no automatic mergeSchema in downstream marts that read shared dimensions. Consumers select explicit columns.
Delta time travel / RESTORE as the rollback path if a bad version slips through.
Balancing a Single Version of Truth with Agility
What works: shared keys, local attributes.
Conform only what's truly cross-domain: Customer, Product, Geography, Date, Org. Use a bus matrix to decide. Everything else stays domain-owned.
When a domain needs extra attributes, it builds an extension table keyed on the conformed SK (e.g. prod_gold.sales.dim_customer_ext) instead of forking the dimension.
Changes to core attributes go through a lightweight PR process on the dimension's repo, not a committee.
This keeps one version of truth for keys and core attributes, while domains move at their own pace.
Common pitfalls:
Each mart building its own "customer" dimension "temporarily". It never stays temporary.
Identity-column surrogate keys that break reproducibility on reloads.
Schema changes pushed without checking lineage first.
Useful docs:
Hope this helps, and I'd like to hear how others handle identity resolution at scale!