<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Re: Photon enabled but a large share of the plan is falling back, cost up and runtime flat in Data Engineering</title>
    <link>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170306#M56278</link>
    <description>&lt;P&gt;Hello,&amp;nbsp;your case I'd test the Python UDFs first: temporarily replace each with an equivalent built-in Spark/SQL expression and compare the Photon time and total runtime. Databricks generally recommends native Spark functions over Python/Scala UDFs for large workloads because of both serialization overhead and engine optimization.&amp;nbsp;I wouldn't use a fixed rule such as “Photon must cover X% of the plan.” The useful comparison is actual runtime and total DBU cost with Photon vs. without it for the same representative 2-TB workload. If Photon only accelerates a small portion while the expensive portion remains in Spark, the DBU multiplier can outweigh the benefit.&lt;/P&gt;</description>
    <pubDate>Thu, 01 Oct 2026 09:05:08 GMT</pubDate>
    <dc:creator>amelia854taylor</dc:creator>
    <dc:date>2026-10-01T09:05:08Z</dc:date>
    <item>
      <title>Photon enabled but a large share of the plan is falling back, cost up and runtime flat</title>
      <link>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170248#M56262</link>
      <description>&lt;P&gt;Hi everyone,&lt;/P&gt;&lt;P&gt;Trying to work out whether this is expected or whether I have misconfigured something.&lt;/P&gt;&lt;P&gt;We enabled Photon on a job cluster running a nightly aggregation over roughly 2TB. The expectation was the usual improvement. What we got instead was runtime essentially unchanged and cost noticeably higher, since Photon carries a DBU multiplier.&lt;/P&gt;&lt;P&gt;Looking at the query profile, a meaningful part of the plan is running outside Photon. I can see the fallback in the plan but I am having trouble turning that observation into an action.&lt;/P&gt;&lt;P&gt;Questions.&lt;/P&gt;&lt;P&gt;What is the most reliable way to identify exactly which operators caused the fallback? I can see that fallback happened, but attributing it to a specific expression in a long query has been guesswork so far.&lt;/P&gt;&lt;P&gt;Are there categories of operation that are known to be unsupported and worth auditing the code for upfront? We use a couple of Python UDFs and some fairly complex nested struct handling, and I suspect one of those is the cause, but I would rather check than guess.&lt;/P&gt;&lt;P&gt;For a plan that partially falls back, is there a rule of thumb for when it is still worth keeping Photon on? At what share of the plan does the multiplier stop paying for itself?&lt;/P&gt;&lt;P&gt;And the practical one: has anyone rewritten a UDF specifically to keep a plan inside Photon, and was the rewrite worth the effort?&lt;/P&gt;&lt;P&gt;Happy to share an anonymised plan if that helps.&lt;/P&gt;&lt;P&gt;Thanks.&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 13:23:46 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170248#M56262</guid>
      <dc:creator>Islam_hoti</dc:creator>
      <dc:date>2026-09-30T13:23:46Z</dc:date>
    </item>
    <item>
      <title>Re: Photon enabled but a large share of the plan is falling back, cost up and runtime flat</title>
      <link>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170285#M56274</link>
      <description>&lt;P&gt;Yeah, this is the classic Photon fallback trap. Flat runtime with higher cost almost always means one unsupported operator is dragging most of your plan back onto Spark, so you're paying the Photon premium on a plan that's barely touching Photon.&lt;/P&gt;
&lt;P&gt;First thing I'd do is stop guessing from the profile colours and run df.explain("formatted") (or EXPLAIN FORMATTED in SQL). Read it top to bottom: anything prefixed Photon... ran in Photon, anything without the prefix ran on Spark, and a ColumnarToRow node marks the spot where Photon gives up and hands off. The explain output also tells you in plain English whether Photon covered the whole query and, if not, which node it choked on.&lt;/P&gt;
&lt;P&gt;I'd put money on it being the Python UDFs, not the structs. The bit that catches people out is that the fallback is contagious. As soon as the plan hits a Python UDF (it shows up as BatchEvalPython), everything downstream of it, your aggregation, the shuffle, any joins, also drops off Photon, not just the UDF itself. So one row-wise UDF poisons the whole back half of the plan, which is why you're seeing a large share fall back rather than a small slice.&lt;/P&gt;
&lt;P&gt;I tested a few variants to be sure:&lt;/P&gt;
&lt;P&gt;- scalar @udf → falls back (BatchEvalPython) and poisons the downstream agg&lt;BR /&gt;- pandas_udf (vectorised) → stays on Photon; it runs as ArrowEvalPython between Photon Arrow nodes and the aggregation carries on in Photon&lt;BR /&gt;- UC SQL Python UDF (CREATE FUNCTION ... LANGUAGE PYTHON) → runs as PhotonScalarUDF, fully supported&lt;BR /&gt;- structs, arrays, explode, higher-order functions like transform → all fine in Photon&lt;/P&gt;
&lt;P&gt;So on the rewrite question: cheapest win is swapping the UDF for native SQL/Spark functions wherever you can. If the logic genuinely has to stay in Python, just moving a scalar @udf to a pandas_udf was enough to keep the surrounding plan on Photon in my testing, and that's far less work than a full rewrite. Registering it as a UC SQL Python UDF works too.&lt;/P&gt;
&lt;P&gt;On whether partial Photon is worth the multiplier, there's no magic number. Photon's DBU rate is roughly 2x (depends on the instance, check the pricing page), so it only pays off if it drops wall-clock to below about half the non-Photon time. I'd just benchmark the job three ways: Photon off, Photon on as-is, and Photon on with the UDF sorted. If you can't get Photon onto the hot path, turning it off for this one job might actually be cheaper. It's not a free win.&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 20:59:08 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170285#M56274</guid>
      <dc:creator>tom_n</dc:creator>
      <dc:date>2026-09-30T20:59:08Z</dc:date>
    </item>
    <item>
      <title>Re: Photon enabled but a large share of the plan is falling back, cost up and runtime flat</title>
      <link>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170306#M56278</link>
      <description>&lt;P&gt;Hello,&amp;nbsp;your case I'd test the Python UDFs first: temporarily replace each with an equivalent built-in Spark/SQL expression and compare the Photon time and total runtime. Databricks generally recommends native Spark functions over Python/Scala UDFs for large workloads because of both serialization overhead and engine optimization.&amp;nbsp;I wouldn't use a fixed rule such as “Photon must cover X% of the plan.” The useful comparison is actual runtime and total DBU cost with Photon vs. without it for the same representative 2-TB workload. If Photon only accelerates a small portion while the expensive portion remains in Spark, the DBU multiplier can outweigh the benefit.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 09:05:08 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170306#M56278</guid>
      <dc:creator>amelia854taylor</dc:creator>
      <dc:date>2026-10-01T09:05:08Z</dc:date>
    </item>
    <item>
      <title>Re: Photon enabled but a large share of the plan is falling back, cost up and runtime flat</title>
      <link>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170310#M56281</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/248581"&gt;@Islam_hoti&lt;/a&gt;&amp;nbsp;,&lt;/P&gt;&lt;P&gt;Photon is fantastic, but it’s an 'all-or-nothing' value proposition. Once you hit a fallback, you lose the vectorized execution benefit for that entire sub-tree, and you're still paying the DBU premium. Here is how we approach auditing this:&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;Identifying the Fallback: The most reliable way to find the culprit isn't just looking at the SQL UI; it’s using the EXPLAIN EXTENDED or EXPLAIN FORMATTED output. Look for the Photon-native flag. If an operator lacks that flag in the plan, it’s your point of fallback.&lt;/LI&gt;&lt;/OL&gt;&lt;UL&gt;&lt;LI&gt;&lt;EM&gt;Tip:&lt;/EM&gt; If you have complex nested struct operations or Python UDFs, check those operators first. Photon is highly optimized for SQL/DataFrame-native expressions, but it is not currently compatible with Python UDFs. If a UDF is in your plan, the entire node—and potentially its parents—will fall back to the standard Spark engine.&lt;/LI&gt;&lt;/UL&gt;&lt;OL&gt;&lt;LI&gt;The Usual Suspects: You’ve already identified them:&lt;/LI&gt;&lt;/OL&gt;&lt;UL&gt;&lt;LI&gt;Python UDFs: These are the #1 cause of Photon fallback. Moving these to Pandas UDFs (Vectorized UDFs) or, ideally, native SQL expressions/built-in functions will almost always resolve the fallback.&lt;/LI&gt;&lt;LI&gt;Complex Nested Types: Photon is improving here, but deeply nested struct or map operations can still trigger fallback if the engine can't map them to a vectorized instruction.&lt;/LI&gt;&lt;/UL&gt;&lt;OL&gt;&lt;LI&gt;The Rule of Thumb: If more than 20-30% of your total execution time (or query plan nodes) is falling back to standard Spark, the DBU multiplier often negates the performance gains. At that point, you’re essentially paying a premium for a partial feature. We’ve found it’s better to be on 'Standard' Spark than to be on 'Photon' with 50% fallback.&lt;/LI&gt;&lt;LI&gt;The 'Rewrite' Effort: Yes, we have rewritten UDFs to stay in Photon, and it has absolutely been worth it. Specifically, rewriting Python UDFs into SQL-native CASE statements or coalesce / concat chains often results in a 3x-5x speedup beyond just 'fixing the fallback'—because you also eliminate the serialization overhead of moving data between the JVM and Python.&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;My recommendation: Audit the query plan specifically for your Python UDFs. Replace them with native Spark/SQL expressions if possible. If you must use UDFs, switch to Vectorized (Pandas) UDFs. If the fallback persists after that, it’s often cheaper to stick to a non-Photon cluster for that specific job.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 10:44:33 GMT</pubDate>
      <guid>https://community.databricks.com/t5/data-engineering/photon-enabled-but-a-large-share-of-the-plan-is-falling-back/m-p/170310#M56281</guid>
      <dc:creator>Khasim_1</dc:creator>
      <dc:date>2026-10-01T10:44:33Z</dc:date>
    </item>
  </channel>
</rss>

