cancel
Showing results forย 
Search instead forย 
Did you mean:ย 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results forย 
Search instead forย 
Did you mean:ย 

Implementing a "Zero-Bus" Architecture in a Lakehouse: Best Practices for Shared Dimensions?

Khasim_1
New Contributor III

Hi Everyone,

I am currently architecting a Medallion Lakehouse and moving toward a "Zero-Bus" (Bus Architecture) approach. My goal is to ensure that our core dimensionsโ€”such as Customer, Product, and Geographyโ€”remain consistent and conformed across all our Gold-layer data marts.

I am interested to hear how the community is tackling this within the Databricks/Unity Catalog ecosystem. Specifically:

  1. Creating the "Golden Record"
  • How are you centralizing the creation of shared dimensions?
  • Are you performing this consolidation in a dedicated processing layer before the Silver stage, or are you utilizing "Shared Silver Dimensions" that all downstream Gold marts consume as a single source of truth?
  1. Governance & Unity Catalog
  • Given the capabilities of Unity Catalog, how are you physically structuring your data?
  • Are you using specific Catalogs or Schemas to isolate conformed dimensions from domain-specific transformation logic?
  1. Change Management & Evolution
  • How do you handle schema evolution for these shared dimensions?
  • What strategies are you using to prevent "breaking changes" from impacting the multiple downstream pipelines that rely on these shared tables?

Iโ€™m particularly curious about how you balance the requirement for a "Single Version of Truth" with the need for agility within the Medallion architecture.

Looking forward to hearing your war stories and best practices!

Data Architect | 13 Years Domain Expertise | Databricks SA Champion Cohort
2 REPLIES 2

LiresaFerizaj
New Contributor III

Great question. Here's what has worked for us.

Conformation: We do entity resolution in Silver (a crosswalk mapping each source system's ID to one enterprise key, with survivorship rules), then publish conformed dimensions once as shared Gold tables that every mart references. For SCD2 at scale, AUTO CDC in Lakeflow Declarative Pipelines handles history and out-of-order data well. We also expose a "current" view, since most consumers only need the latest version.

Unity Catalog: Conformed dimensions live in a dedicated catalog owned by a central team, with read-only access for marts and PII protected via column masks. Domain teams can extend but not modify: they build their own extension tables or rollups keyed off the conformed key, so nobody is blocked.

Schema evolution: Treat views as the contract. Additive changes go in place; breaking changes ship as a new versioned view with a deprecation window. UC lineage shows downstream impact before any change. For contract testing, we use pipeline expectations on the provider side (unique keys, one current row per business key) and schema checks in CI on the consumer side.

Lessons: Only conform what's truly shared, start with the dimensions causing the most reconciliation pain (usually Customer), and give each one a clear business owner. Most conformity issues turn out to be definition disputes, not technical ones.

Liresa Ferizaj

Islam_hoti
New Contributor III

Hi,

Great question. This is where most Medallion implementations either scale or get messy. Here's the pattern that has worked well, based on Kimball conformed dimensions mapped onto Unity Catalog.

Quick note on naming: in Databricks, "Zerobus" is a product (Zerobus Ingest, a streaming ingestion API in Lakeflow Connect). Calling this the Kimball Bus Architecture avoids confusion in searches and discussions.

1. Creating the Golden Record

Don't add a separate layer before Silver. Do the conforming in Silver, as its own zone:

Bronze: raw data, one table set per source system (CRM, ERP, e-commerceโ€ฆ).
Silver (per source): cleansed and typed, still one set per source.
Silver/Conformed: where the golden record gets built. This step does:
Identity resolution: a crosswalk (key map) table, (source_system, source_key) โ†’ customer_sk.
Survivorship rules: which source wins for which attribute. Versioned in Git.
Surrogate keys: deterministic hash keys (e.g. xxhash64 or sha2 of the business key) are reproducible and avoid concurrency problems with identity columns.
SCD history: Lakeflow Declarative Pipelines' AUTO CDC API handles SCD Type 1/2 and out-of-order events for you.
Gold: domain marts that consume the conformed dimensions and never re-derive them.

The key rule is one writer per conformed dimension: one pipeline, one owning team, one table. Every Gold mart references it and never copies or rebuilds it.

If you already run a real MDM tool (Reltio, Informatica, Profisee, etc.), treat its output as a source. Don't rebuild MDM in Spark.

Also add unknown / inferred member rows (e.g. sk = -1) so facts that arrive before their dimension rows still load and don't get dropped.

2. Governance & Unity Catalog Structure

A layout that isolates shared dimensions from domain logic:

prod_silver.<source>.* -- per-source cleansed data
prod_conformed.dims.* -- physical conformed dims (platform team writes)
prod_conformed.published.* -- views = the public contract
prod_gold.sales.* -- domain marts (sales team)
prod_gold.finance.* -- domain marts (finance team)
A dedicated catalog or schema for conformed dimensions. Only the owning team can write to it. Domain teams get SELECT on the published views only.
Views as the contract layer. Consumers query published.dim_customer, a view over the physical table. You can then change physical storage without breaking anyone.
Central security. Put row filters and column masks on the conformed dimensions once, and every mart inherits the PII protection.
Catalogs per environment (dev/test/prod), with workspaceโ€“catalog binding so dev jobs can't touch prod.
Tags for owner, classification and deprecation status.
Lineage (system.access.table_lineage / column_lineage) to see exactly who consumes each dimension. You'll need this for change management.
Delta Sharing if marts live in another metastore or region.
3. Change Management & Evolution

Treat each conformed dimension as a data product with a contract: grain, keys, schema, SLA and owner, with the DDL in Git, deployed through CI/CD (Asset Bundles or Terraform).

Non-breaking changes (allowed anytime):

Adding nullable columns.
Widening types (int โ†’ bigint, etc.) via delta.enableTypeWidening, without rewriting data.

Breaking changes (renames, drops, grain or key changes):

Build dim_customer_v2 next to v1.
Keep published.dim_customer pointing to v1, and expose published.dim_customer_v2.
Use lineage to find all downstream consumers and notify their owners.
Tag v1 as deprecated with a sunset date. Migrate consumers, then switch the view and drop v1.

Column mapping (delta.columnMapping.mode = 'name') lets you rename or drop physically without a rewrite. For consumers, though, a rename is still a breaking change, so do it behind the view.

Guardrails that prevent incidents:

Writeโ€“Auditโ€“Publish: build into a staging table and run expectations (unique SK, non-null business key, exactly one current row per key in SCD2). Publish only if they pass.
No SELECT * and no automatic mergeSchema in downstream marts that read shared dimensions. Consumers select explicit columns.
Delta time travel / RESTORE as the rollback path if a bad version slips through.
Balancing a Single Version of Truth with Agility

What works: shared keys, local attributes.

Conform only what's truly cross-domain: Customer, Product, Geography, Date, Org. Use a bus matrix to decide. Everything else stays domain-owned.
When a domain needs extra attributes, it builds an extension table keyed on the conformed SK (e.g. prod_gold.sales.dim_customer_ext) instead of forking the dimension.
Changes to core attributes go through a lightweight PR process on the dimension's repo, not a committee.

This keeps one version of truth for keys and core attributes, while domains move at their own pace.

Common pitfalls:

Each mart building its own "customer" dimension "temporarily". It never stays temporary.
Identity-column surrogate keys that break reproducibility on reloads.
Schema changes pushed without checking lineage first.

Useful docs:

Hope this helps, and I'd like to hear how others handle identity resolution at scale!