Sanjeeb2024
Valued Contributor

Thanks for the question. This is a very important and interesting question. Please find details from my side.

  1. How are you using File Arrival Triggers and Table Update Triggers to avoid running compute against empty or unchanged sources? Have you seen a measurable DBU savings from this switch?                                              [San] The File Arrival Trigger and Table update Triggers are controlled by control plane orchestration and there is no DBU cost, DBUs start when Databricks actually provisions and runs the workload. However there is a cost associated with this how you configure the triggers ( depending upon the cloud provider, but its should be minimal ideally).
  2. For cross-workspace dependencies (Job A in Workspace 1 needs to complete before Job B in Workspace 2 starts), are you using a "Signal Table" pattern in Unity Catalog, or is there a more native way to handle this in Lakeflow?                                                                                                                                                                   [ San:] As far I know there is no out of the box solution exists and this is something Databricks can think of down the line, at present Lakeflow task dependencies are designed for tasks within the same job/workspace. However you can create some control table approach or create a zero byte file ( flag file) approach to signal and put the dependency between jobs. This is just a work around but it will work.
  3. How do you handle "fan-out" dependencies—i.e., one upstream table update needs to trigger 5-6 downstream jobs owned by different teams? Are you managing this centrally, or letting each team own their own trigger subscription?                                                                                                                                                             [ San:] It depends. If you want to give the autonomy to the downstream systems, better to create flag file and provide the signal about the process completion. When Databricks will come up the dependency between workspaces job, possibly a direct dependency can be established.

Lets hear from other Databricks experts view on this. I believer Airflow is a good orchestration tool and Databricks needs to improve certain areas before migrate complete job orchestration to Databricks eco system ( This is my personal opinion).

Regards - Sanjeeb

Sanjeeb Mohapatra