<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Intialization stage is taking time in lakehouse pipeline in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/intialization-stage-is-taking-time-in-lakehouse-pipeline/m-p/161485#M55027</link>
    <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/124640"&gt;@Anish_2&lt;/a&gt;,&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea" data-pm-slice="1 1 []"&gt;The message...&amp;nbsp;Reported flow time metrics for flowName: 'pipelines.flowTimeMetrics.missingFlowName'..is usually not enough on its own to identify the root cause. I would treat it as a metrics/observability signal rather than the primary failure, and I would focus on the surrounding pipeline events for the same update. The public docs describe the Lakeflow event log as the main source for execution progress, data quality, lineage, and streaming metrics, and they recommend using it to inspect pipeline behaviour in detail.&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;A good next step is to query the event log for the same update_id and look at the events immediately before and after these messages. Databricks documents the event_log('&amp;lt;pipeline-id&amp;gt;') table-valued function for this purpose in &lt;A href="https://docs.databricks.com/aws/en/ldp/monitor-event-logs" rel="noopener noreferrer nofollow" target="_blank"&gt;Pipeline event log&lt;/A&gt; and &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/observability" rel="noopener noreferrer nofollow" target="_blank"&gt;Monitor pipelines&lt;/A&gt;.&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;For example, you can review the latest flow progress events like this:&lt;/P&gt;
&lt;PRE class="_1ibi0s3d1" dir="auto"&gt;&lt;CODE class="language-sql"&gt;SELECT timestamp, level, event_type, message, details, origin
FROM event_log('&amp;lt;pipeline-id&amp;gt;')
WHERE origin.update_id = '&amp;lt;update-id&amp;gt;'
ORDER BY timestamp;&lt;/CODE&gt;&lt;/PRE&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;When reviewing that output, I would specifically check for these patterns:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;WAITING_FOR_RESOURCES, autoscaling, or cluster resize activity, which can indicate the pipeline is spending time acquiring or resizing compute rather than actually processing data. The autoscaling docs note that scaling behaviour and worker limits can affect latency, and that classic pipelines expose autoscaling events in the event log.&lt;/LI&gt;
&lt;LI&gt;SETTING_UP_TABLES, STARTING, RUNNING, stream_progress, or operation_progress events, which help confirm whether the pipeline is making progress but is simply slow to initialize. Databricks documents these monitoring surfaces in &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/monitoring-ui" rel="noopener noreferrer nofollow" target="_blank"&gt;Monitor pipelines in the UI&lt;/A&gt; and &lt;A href="https://docs.databricks.com/aws/en/ldp/monitor-event-logs" rel="noopener noreferrer nofollow" target="_blank"&gt;Pipeline event log&lt;/A&gt;.&lt;/LI&gt;
&lt;LI&gt;Any explicit error events in the same update, especially driver, source, schema, or permission-related failures. The monitoring UI and event log are the recommended places to inspect these details.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If the pipeline is using classic compute, a 5 - 6 minute initialisation can sometimes be explained by compute startup or autoscaling behaviour rather than by the missingFlowName message itself. Databricks recommends &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/auto-scaling" rel="noopener noreferrer nofollow" target="_blank"&gt;Enhanced autoscaling&lt;/A&gt; for Lakeflow pipelines and explains that compute configuration can influence latency and startup behaviour.&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;If this is a new pipeline or you are re-evaluating the compute setup, it is also worth reviewing &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/best-practices" rel="noopener noreferrer nofollow" target="_blank"&gt;Best practices for Lakeflow Spark Declarative Pipelines&lt;/A&gt;, which recommends using the event log for observability and discusses compute choices such as serverless and autoscaling.&lt;/P&gt;
&lt;P class="p1"&gt;&lt;FONT size="2" color="#FF6600"&gt;&lt;STRONG&gt;&lt;I&gt;If this answer resolves your question, could you mark it as “Accept as Solution”? That helps other users quickly find the correct fix.&lt;/I&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;I&gt;&lt;/I&gt;&lt;/P&gt;</description>
    <pubDate>Sun, 05 Jul 2026 13:51:01 GMT</pubDate>
    <dc:creator>Ashwin_DSA</dc:creator>
    <dc:date>2026-07-05T13:51:01Z</dc:date>
    <item>
      <title>Intialization stage is taking time in lakehouse pipeline</title>
      <link>https://community.databricks.com/t5/data-engineering/intialization-stage-is-taking-time-in-lakehouse-pipeline/m-p/161333#M55019</link>
      <description>&lt;P&gt;Hello Team,&lt;/P&gt;&lt;P&gt;My intialization stage in lakehouse pipeline is taking 5-6 mins. When i checked event log table,below are stats&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="Anish_2_0-1783115921054.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/28545iBF71C7E501748C62/image-size/medium?v=v2&amp;amp;px=400" role="button" title="Anish_2_0-1783115921054.png" alt="Anish_2_0-1783115921054.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;Message&amp;nbsp;&lt;SPAN&gt;&lt;STRONG&gt;Reported flow time metrics for flowName: 'pipelines.flowTimeMetrics.missingFlowName'.&lt;/STRONG&gt; is repeatedly coming for 4-5 mins. Can someone help what can be possible root cause for this message?&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 03 Jul 2026 22:01:20 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/intialization-stage-is-taking-time-in-lakehouse-pipeline/m-p/161333#M55019</guid>
      <dc:creator>Anish_2</dc:creator>
      <dc:date>2026-07-03T22:01:20Z</dc:date>
    </item>
    <item>
      <title>Re: Intialization stage is taking time in lakehouse pipeline</title>
      <link>https://community.databricks.com/t5/data-engineering/intialization-stage-is-taking-time-in-lakehouse-pipeline/m-p/161485#M55027</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/124640"&gt;@Anish_2&lt;/a&gt;,&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea" data-pm-slice="1 1 []"&gt;The message...&amp;nbsp;Reported flow time metrics for flowName: 'pipelines.flowTimeMetrics.missingFlowName'..is usually not enough on its own to identify the root cause. I would treat it as a metrics/observability signal rather than the primary failure, and I would focus on the surrounding pipeline events for the same update. The public docs describe the Lakeflow event log as the main source for execution progress, data quality, lineage, and streaming metrics, and they recommend using it to inspect pipeline behaviour in detail.&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;A good next step is to query the event log for the same update_id and look at the events immediately before and after these messages. Databricks documents the event_log('&amp;lt;pipeline-id&amp;gt;') table-valued function for this purpose in &lt;A href="https://docs.databricks.com/aws/en/ldp/monitor-event-logs" rel="noopener noreferrer nofollow" target="_blank"&gt;Pipeline event log&lt;/A&gt; and &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/observability" rel="noopener noreferrer nofollow" target="_blank"&gt;Monitor pipelines&lt;/A&gt;.&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;For example, you can review the latest flow progress events like this:&lt;/P&gt;
&lt;PRE class="_1ibi0s3d1" dir="auto"&gt;&lt;CODE class="language-sql"&gt;SELECT timestamp, level, event_type, message, details, origin
FROM event_log('&amp;lt;pipeline-id&amp;gt;')
WHERE origin.update_id = '&amp;lt;update-id&amp;gt;'
ORDER BY timestamp;&lt;/CODE&gt;&lt;/PRE&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;When reviewing that output, I would specifically check for these patterns:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;WAITING_FOR_RESOURCES, autoscaling, or cluster resize activity, which can indicate the pipeline is spending time acquiring or resizing compute rather than actually processing data. The autoscaling docs note that scaling behaviour and worker limits can affect latency, and that classic pipelines expose autoscaling events in the event log.&lt;/LI&gt;
&lt;LI&gt;SETTING_UP_TABLES, STARTING, RUNNING, stream_progress, or operation_progress events, which help confirm whether the pipeline is making progress but is simply slow to initialize. Databricks documents these monitoring surfaces in &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/monitoring-ui" rel="noopener noreferrer nofollow" target="_blank"&gt;Monitor pipelines in the UI&lt;/A&gt; and &lt;A href="https://docs.databricks.com/aws/en/ldp/monitor-event-logs" rel="noopener noreferrer nofollow" target="_blank"&gt;Pipeline event log&lt;/A&gt;.&lt;/LI&gt;
&lt;LI&gt;Any explicit error events in the same update, especially driver, source, schema, or permission-related failures. The monitoring UI and event log are the recommended places to inspect these details.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;If the pipeline is using classic compute, a 5 - 6 minute initialisation can sometimes be explained by compute startup or autoscaling behaviour rather than by the missingFlowName message itself. Databricks recommends &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/auto-scaling" rel="noopener noreferrer nofollow" target="_blank"&gt;Enhanced autoscaling&lt;/A&gt; for Lakeflow pipelines and explains that compute configuration can influence latency and startup behaviour.&lt;/P&gt;
&lt;P class="wnfdntt _1ibi0s3f5 _1ibi0s3ce _1ibi0s3ea"&gt;If this is a new pipeline or you are re-evaluating the compute setup, it is also worth reviewing &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ldp/best-practices" rel="noopener noreferrer nofollow" target="_blank"&gt;Best practices for Lakeflow Spark Declarative Pipelines&lt;/A&gt;, which recommends using the event log for observability and discusses compute choices such as serverless and autoscaling.&lt;/P&gt;
&lt;P class="p1"&gt;&lt;FONT size="2" color="#FF6600"&gt;&lt;STRONG&gt;&lt;I&gt;If this answer resolves your question, could you mark it as “Accept as Solution”? That helps other users quickly find the correct fix.&lt;/I&gt;&lt;/STRONG&gt;&lt;/FONT&gt;&lt;I&gt;&lt;/I&gt;&lt;/P&gt;</description>
      <pubDate>Sun, 05 Jul 2026 13:51:01 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/intialization-stage-is-taking-time-in-lakehouse-pipeline/m-p/161485#M55027</guid>
      <dc:creator>Ashwin_DSA</dc:creator>
      <dc:date>2026-07-05T13:51:01Z</dc:date>
    </item>
  </channel>
</rss>

