- Mark as New
- Bookmark
- Subscribe
- Mute
- Subscribe to RSS Feed
- Permalink
- Report Inappropriate Content
02-20-2024 07:15 AM - edited 02-20-2024 07:17 AM
I imagine Databricks would have to alter the schema of their Jobs API to implement a solution where schedule would also be an Id field instead of just Job Id. I imagine it would be possible to have a lookup table and append what parameters were run and then infer what the next parameter would be, but that would increase ETL time.
Our team haven't seen the need to implement more complicated workflows, here all our workflows have 1 task and that is to run a notebook. That one notebook runs different endpoints/logic/methods using a parallelism/async logic so that is our way of implementing multiple "tasks".
We build solutions where ETL time is an important factor, here multiple tasks also create an issue. For example if you create a Task1 -> Task2 -> Task3 that does a simple print(1) you will see that there is an overhead of approximately 7 seconds between tasks.