cancel
Showing results forĀ 
Search instead forĀ 
Did you mean:Ā 
Data Engineering
Join discussions on data engineering best practices, architectures, and optimization strategies within the Databricks Community. Exchange insights and solutions with fellow data engineers.
cancel
Showing results forĀ 
Search instead forĀ 
Did you mean:Ā 

Using Databricks Asset Bundles and Lakeflow Jobs in a Real Project

gowri_databrick
New Contributor III

Hi everyone,

I’m working through Databricks deployment and orchestration concepts and wanted to understand how Databricks Asset Bundles and Lakeflow Jobs fit together.

Real-world scenario:
Imagine an e-commerce company has a customer data pipeline that processes new data every night. The project contains notebooks, pipeline code, and a job that needs to run on a schedule.

In this situation:

  • How would Databricks Asset Bundles help manage and deploy the project resources?
  • How would Lakeflow Jobs be used to schedule and run the pipeline?
  • How do these two concepts work together as part of a deployment workflow?

I’m especially interested in understanding how this would be handled in a real data engineering project.

Thanks!

4 REPLIES 4

Satyasai
New Contributor III

Please read this below Article, it was mentioned clearly

https://kaninipro.com/2025/11/30/deploying-lakeflow-jobs-with-databricks-asset-bundles/

 

srini_ve
Contributor

@gowri_databrick 

A simple way to understand this is:

Asset Bundles are used to deploy the project, while Lakeflow Jobs are used to run the deployed pipeline.

Practical scenario

Imagine an e-commerce company has a customer pipeline that runs every night.

The project contains:

Customer ingestion notebook
Customer transformation code
A Lakeflow Job to run these steps


Asset Bundles:

The developer keeps the notebooks, code and job configuration in Git.

For example:

Developer → Git → Databricks Asset Bundle

The Asset Bundle describes what needs to be deployed, such as notebooks, jobs, tasks and configuration.

When the code is ready, the bundle can be deployed to DEV, TEST and PROD without manually creating the same resources in each environment.

So, Asset Bundles mainly help with version control and consistent deployment.


Lakeflow Jobs:


Once the job is deployed, Lakeflow Jobs is responsible for running the pipeline.

For example, we can configure the job to run every night at 1 AM.

The job could have tasks like:

Ingest customer data → Transform customer data → Load the final table

Lakeflow Jobs manages the task order, dependencies, scheduling, retries and execution.


Asset Bundle + Lakeflow


In a real project, the flow would be something like:

Step 1: Developer changes the customer pipeline code.

Step 2: The changes are committed to Git.

Step 3: The Asset Bundle is deployed through the CI/CD pipeline.

Step 4: The bundle creates or updates the Lakeflow Job in the target environment.

Step 5: Lakeflow Jobs runs the pipeline based on the configured schedule.

So, I normally think of it this way:

Asset Bundles :- Deploy and manage the project

Lakeflow Jobs :- Schedule and run the project

They work together nicely because the job itself can be defined as code inside the Asset Bundle, rather than manually creating and configuring jobs separately in each environment.

balajij8
Esteemed Contributor II

@gowri_databrick 

Automation Bundles serve as the infrastructure as code layer for your e-commerce customer data pipeline. You define everything in databricks.yml file with resource definitions in resources/*.yml - notebooks in src/, pipeline configurations, job schedules and Unity Catalog schemas and volumes. The advantage is multi environment targeting - you can define variables like catalog, schema, and warehouse_id once, then override them per target (dev uses dev_catalog/dev_schema, prod uses prod_catalog/prod_schema). A single databricks bundle validate catches configuration errors and databricks bundle deploy pushes all resources to the workspace mapped atomically. It eliminates manual deployment and makes your pipeline reproducible, version-controlled and CI/CD based.

Lakeflow Jobs provide the orchestration engine that actually executes your nightly pipeline. In your e-commerce scenario, you'd define a multi-task DAG - an extract task pulls new customer data from source systems, a transform task (with depends_on - extract) cleans and enriches it, and a load task writes results to Delta tables - each task using run_if - ALL_SUCCESS to chain execution. The job runs on a cron schedule (0 0 2 * * ? for 2 AM nightly) and can use job clusters, autoscaling clusters or serverless compute (by omitting cluster config for notebook/Python tasks). Task types include notebooks, Python scripts, SQL queries and even pipeline triggers for Spark Declarative Pipelines. You can parameterize tasks with job-level variables (date: "{{start_date}}") accessed via dbutils.widgets.get() in notebooks and set permissions for different teams.

They work together seamlessly - DABs defines the job as a resource in YAML and Lakeflow Jobs executes it.

Coffee77
Honored Contributor III

Here is how I am using DAB in a real project. I hope it helps.


Lifelong Solution Architect Learner | Coffee & Data