cancel
Showing results for 
Search instead for 
Did you mean: 
Administration & Architecture
Explore discussions on Databricks administration, deployment strategies, and architectural best practices. Connect with administrators and architects to optimize your Databricks environment for performance, scalability, and security.
cancel
Showing results for 
Search instead for 
Did you mean: 

Deployment Jobs structured yaml.

AnhPT
New Contributor III

 

  1. Does DAB support any form of task reuse or shared task templates across job definitions — so the wrapper tasks don't need to be physically repeated in every file?

  2. Is there a recommended pattern to externalize the orchestration wrapper (logging, skip logic) from individual job definitions — perhaps as a shared job that other jobs trigger?

  3. For teams using code generation to produce DAB YAML at scale: how do you handle jobs that deviate from the standard template without forking into unmaintained hand-written files?



2 REPLIES 2

AnhPT
New Contributor III

1.Current State
The deployment uses Dâ Asset Bundle (DAB) with the following structure:

entries/ ← developer-authored business config
---------*.yml ← what tables to read/write, pre-check conditions
resources/ ← dbx Bundle YAML, ready to deploy 

Every generated resources/*.yml contains an identical 6-task wrapper pattern:

- task_key: set_params # identical in all 
- task_key: create_log # identical in all 
- task_key: skip_run # identical in all 
- task_key: finish_log_skipped # identical in all
- task_key: finish_log_success # identical in all

← full job definitions consumed by `databricks bundle deploy`
resources/common/ ← shared variables: clusters, tags, notifications
joblib/helper/entry_converter/templates/
------ single_job_resource_template.j2
------ multi_jobs_resource_template.j2

juan_maedo
Contributor

Hi @AnhPT ,

1 and 2. I’ll combine these two since it might be helpful—I can tell you that I’ve used this approach successfully. In a situation I’ve encountered with similar archetypes—where utilities or steps were exactly the same across bundles—and where, most importantly, if one changes, they all must change (cluster configurations, catalogs, run and manage permissions, etc.)—I’ve used a cross-bundle shared repository to store those utilities or core components used by the rest. This was accessed by defining the relative paths to it. This way, if I update a function or a notebook, when I deploy it to the next environment, every new job will use it. In my case, I deploy it using SPNs in a folder with read-only access for users, while the SPNs have execute permissions to compile it in the jobs—ensuring that “no one” can make manual changes.

 

3. Without a doubt, the best option for this—and to streamline the process (including agents)—is to use custom templates during initialization, whether from a Git/DevOps tool or directly within the workspaces. https://docs.databricks.com/aws/en/dev-tools/bundles/templates#custom-bundle-templates