Comment
Databricks Employee
Databricks Employee

Hi @antoalphi ,

Thanks for the feedback!

The choice between Streaming Tables (ST) and Materialized Views (MV) really comes down to your required processing semantics:

  • If a row from the source needs to be processed only once / the logic never looks back at historical data -> go for a Streaming Table
  • If your logic requires complex aggregations, joins, or updates to existing records ->  you have to go for a Materialized View

For example, Streaming Table is typically not suitable for aggregations and joins, just because the semantics need to look back to existing records in the table in order to perform those. There are exceptions which I plan to discuss in my next post.

If you need to ingest thousands of sources into UC tables in an incremental way with lightweight row-by-row processing - Streaming Tables is a good choice.

If you have to create aggregations / joins on those tables - you have to go with Materialized Views on top. 

All in all, wherever you have thousands of tables or just few of them - the choice really depends on the business logic.

A note regarding large-scale environments:

  • One pipeline can contain up to 1000 datasets (including Materialized Views, Streaming Tables, and Temporary Views). At a large scale, pipeline organization is critical since all datasets within a single pipeline share the same compute resources and lifecycle. If one table hangs, it can impact the entire group.

If you have a specific pipeline diagram or a list of requirements, feel free to share it. I’d be happy to provide a more tailored vision for the setup!