<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Best practices for SLA monitoring and automated retries across hundreds of Lakeflow Jobs in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168435#M55921</link>
    <description>&lt;P&gt;Hi everyone,&lt;/P&gt;&lt;P&gt;We are in the process of migrating our legacy Airflow DAGs (largely time-based cron schedules) to&amp;nbsp;Lakeflow Jobs, and I want to fully leverage the platform's event-driven capabilities rather than just replicating the old "timer-based" pattern.&lt;/P&gt;&lt;P&gt;A few questions for those further along in this migration:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;How are you using&amp;nbsp;File Arrival Triggers&amp;nbsp;and&amp;nbsp;Table Update Triggers&amp;nbsp;to avoid running compute against empty or unchanged sources? Have you seen a measurable DBU savings from this switch?&lt;/LI&gt;&lt;LI&gt;For cross-workspace dependencies (Job A in Workspace 1 needs to complete before Job B in Workspace 2 starts), are you using a "Signal Table" pattern in Unity Catalog, or is there a more native way to handle this in Lakeflow?&lt;/LI&gt;&lt;LI&gt;How do you handle&amp;nbsp;"fan-out" dependencies—i.e., one upstream table update needs to trigger 5-6 downstream jobs owned by different teams? Are you managing this centrally, or letting each team own their own trigger subscription?&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Would love to hear the "messy" lessons learned, not just the documented happy path.&lt;/P&gt;</description>
    <pubDate>Sat, 12 Sep 2026 18:36:32 GMT</pubDate>
    <dc:creator>Khasim_1</dc:creator>
    <dc:date>2026-09-12T18:36:32Z</dc:date>
    <item>
      <title>Best practices for SLA monitoring and automated retries across hundreds of Lakeflow Jobs</title>
      <link>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168435#M55921</link>
      <description>&lt;P&gt;Hi everyone,&lt;/P&gt;&lt;P&gt;We are in the process of migrating our legacy Airflow DAGs (largely time-based cron schedules) to&amp;nbsp;Lakeflow Jobs, and I want to fully leverage the platform's event-driven capabilities rather than just replicating the old "timer-based" pattern.&lt;/P&gt;&lt;P&gt;A few questions for those further along in this migration:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;How are you using&amp;nbsp;File Arrival Triggers&amp;nbsp;and&amp;nbsp;Table Update Triggers&amp;nbsp;to avoid running compute against empty or unchanged sources? Have you seen a measurable DBU savings from this switch?&lt;/LI&gt;&lt;LI&gt;For cross-workspace dependencies (Job A in Workspace 1 needs to complete before Job B in Workspace 2 starts), are you using a "Signal Table" pattern in Unity Catalog, or is there a more native way to handle this in Lakeflow?&lt;/LI&gt;&lt;LI&gt;How do you handle&amp;nbsp;"fan-out" dependencies—i.e., one upstream table update needs to trigger 5-6 downstream jobs owned by different teams? Are you managing this centrally, or letting each team own their own trigger subscription?&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;Would love to hear the "messy" lessons learned, not just the documented happy path.&lt;/P&gt;</description>
      <pubDate>Sat, 12 Sep 2026 18:36:32 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168435#M55921</guid>
      <dc:creator>Khasim_1</dc:creator>
      <dc:date>2026-09-12T18:36:32Z</dc:date>
    </item>
    <item>
      <title>Re: Best practices for SLA monitoring and automated retries across hundreds of Lakeflow Jobs</title>
      <link>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168460#M55924</link>
      <description>&lt;P&gt;Just answer the details in another thread. Please check.&lt;/P&gt;</description>
      <pubDate>Sun, 13 Sep 2026 11:55:06 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168460#M55924</guid>
      <dc:creator>Sanjeeb2024</dc:creator>
      <dc:date>2026-09-13T11:55:06Z</dc:date>
    </item>
    <item>
      <title>Re: Best practices for SLA monitoring and automated retries across hundreds of Lakeflow Jobs</title>
      <link>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168800#M55987</link>
      <description>&lt;P&gt;Hi,&lt;/P&gt;&lt;P&gt;Migrating the same way. The messy parts.&lt;/P&gt;&lt;P&gt;On file arrival triggers, first check whether your external locations have file events enabled, because the behaviour differs sharply. Without them you get a cap of 50 triggered jobs per workspace and at most 10,000 files in the monitored path. With file events, the file limit goes away.&lt;/P&gt;&lt;P&gt;Four gotchas that cost us time. Overwriting a file with the same name does not fire the trigger, so any producer writing to a stable filename will silently never run. Conversely, modifying an old file whose metadata has aged out of the internal retention window is treated as a new arrival and fires a spurious run. With file events on, a trigger set on a subpath can time out and error when there is heavy unrelated churn elsewhere under the same root, and the fix is to create a volume mapped to exactly the directory you want and trigger on its root. Worst of all, on S3 and GCS a trigger pointed at a deleted or non existent path never errors and never fires, so you get silence rather than an alert.&lt;/P&gt;&lt;P&gt;Also use the two rate options. Minimum time between triggers is a cooldown, wait after last change is a debounce that resets on each arrival. The second one is what stops a job firing on the first file of a batch of 200.&lt;/P&gt;&lt;P&gt;On savings, frame it carefully. The triggers add no cost beyond cloud listing charges, so the saving is entirely the compute you were spending on runs that found nothing. Measure your current empty run rate before promising a number.&lt;/P&gt;&lt;P&gt;On question two, you likely do not need a signal table. Unity Catalog objects live at the metastore rather than the workspace, so a table update trigger in workspace two can watch a table written from workspace one as long as both share a metastore. The table is the signal. Worth testing how it reacts to optimize and vacuum commits on the watched table, since those are writes too.&lt;/P&gt;&lt;P&gt;On question three, table update triggers are a subscription model, so fan out falls out naturally and the producer does nothing. The messy part is what you lose from Airflow: no single place to see the whole graph. Producers have no idea who depends on them until something breaks, and lineage is the only thing filling that gap.&lt;/P&gt;</description>
      <pubDate>Wed, 16 Sep 2026 11:54:40 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/best-practices-for-sla-monitoring-and-automated-retries-across/m-p/168800#M55987</guid>
      <dc:creator>Islam_hoti</dc:creator>
      <dc:date>2026-09-16T11:54:40Z</dc:date>
    </item>
  </channel>
</rss>

