With Apache Spark™ Declarative Pipelines (SDP) you describe the tables you want, and the engine works out how to build them, tracks the dependencies, and keeps the data fresh. Define streaming tables and materialized views, enforce data quality with expectations, and let SDP handle incremental processing, orchestration, and recovery.
The catalog has two SDP courses on the Data Engineer track, both with a free self-paced option. Climb from Associate to Professional. 👇
1 Build the foundations · Associate
Data Engineer · Associate
Build Data Pipelines with Apache Spark Declarative Pipelines
The essential concepts and skills to build pipelines with SDP for incremental batch or streaming ingestion through streaming tables and materialized views. Develop and debug ETL in the multi-file editor with SQL (Python examples provided), see how the pipeline graph tracks data dependencies, and configure compute, trigger modes, and advanced options. Add data quality expectations to validate and enforce integrity, put pipelines into production with scheduling and event logging, and implement Change Data Capture using AUTO CDC INTO for SCD Type 1 and Type 2.
2 Master advanced techniques · Professional
Ready for production scale? Go deep on design patterns, cross-platform integration, and bulletproof data quality. Part of the Advanced Data Engineering with Databricks series.
Data Engineer · Professional
Advanced Techniques with Apache Spark Declarative Pipelines
Advanced design patterns for production-grade streaming pipelines. Build multi-flow pipelines that ingest multi-source data into a unified Bronze table, apply Liquid Clustering and data quality expectations across the Silver and Gold layers, and implement the Multiplex Streaming pattern with Iceberg UniForm for cross-platform access. Automate SCD Type 2 history tracking with AUTO CDC INTO, handle schema evolution, and design zero-data-loss quarantine pipelines to audit and manage invalid records.
Instructor-led available in English, 日本語, Português BR, and 한국어.