<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>All Generative AI posts</title>
    <link>https://community.databricks.com/t5/generative-ai/bd-p/GenAI-Insight-Hub</link>
    <description>All Generative AI posts</description>
    <pubDate>Fri, 02 Oct 2026 00:40:38 GMT</pubDate>
    <dc:creator>GenAI-Insight-Hub</dc:creator>
    <dc:date>2026-10-02T00:40:38Z</dc:date>
    <item>
      <title>Re: Genie Agent SDK visualization</title>
      <link>https://community.databricks.com/t5/generative-ai/genie-agent-sdk-visualization/m-p/170361#M2138</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/170177"&gt;@haaland&lt;/a&gt;, I took a look at both internal and external documentation and here is what I found.&lt;/P&gt;
&lt;P&gt;Glad it's working, and &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/135464"&gt;@Gecofer&lt;/a&gt; and &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/249473"&gt;@ivanvyd&lt;/a&gt; had the right diagnosis. &lt;CODE&gt;enable_visualization&lt;/CODE&gt; was never removed. It shipped in &lt;CODE&gt;databricks-sdk&lt;/CODE&gt; v0.120.0 (July 2026), alongside the &lt;CODE&gt;viz&lt;/CODE&gt; attachment type and &lt;CODE&gt;download_message_attachment_visualization()&lt;/CODE&gt;, and nothing through 0.141.0 took it back out. So a 0.141.0 version string next to a signature that lacks the argument means the files on disk and the code running in your Python process had drifted apart: either a stale process still holding the classes your &lt;CODE&gt;WorkspaceClient&lt;/CODE&gt; was built from, or a second interpreter or SDK copy doing the loading.&lt;/P&gt;
&lt;P&gt;The checks Gecofer and Ivan suggested are the way to confirm which:&lt;/P&gt;
&lt;PRE&gt;&lt;CODE class="language-python"&gt;import inspect, sys
import databricks.sdk

print(sys.executable)
print(databricks.sdk.__version__)
print(databricks.sdk.__file__)
print(inspect.signature(w.genie.start_conversation_and_wait))
&lt;/CODE&gt;&lt;/PRE&gt;
&lt;P&gt;In a notebook, upgrade, restart Python, then rebuild the client:&lt;/P&gt;
&lt;PRE&gt;&lt;CODE class="language-python"&gt;%pip install --upgrade databricks-sdk
dbutils.library.restartPython()
&lt;/CODE&gt;&lt;/PRE&gt;
&lt;P&gt;After the restart the signature should show &lt;CODE&gt;enable_visualization: Optional[bool] = None&lt;/CODE&gt;. If it still doesn't, compare &lt;CODE&gt;sys.executable&lt;/CODE&gt; and &lt;CODE&gt;databricks.sdk.__file__&lt;/CODE&gt; against the environment where you ran the install. For a Databricks App, pin &lt;CODE&gt;databricks-sdk&amp;gt;=0.120.0&lt;/CODE&gt; in &lt;CODE&gt;requirements.txt&lt;/CODE&gt; and redeploy.&lt;/P&gt;
&lt;P&gt;On what the flag actually does: it asks Genie to generate a visualization but doesn't guarantee one. When the question and data support a chart, the response carries a &lt;CODE&gt;viz&lt;/CODE&gt; attachment (Beta) next to the query attachment, with a &lt;CODE&gt;title&lt;/CODE&gt; and the &lt;CODE&gt;query_attachment_id&lt;/CODE&gt; it came from. For a rendered PNG, call &lt;CODE&gt;w.genie.download_message_attachment_visualization(name=...)&lt;/CODE&gt; with &lt;CODE&gt;name&lt;/CODE&gt; in the form &lt;CODE&gt;spaces/{space_id}/conversations/{conversation_id}/messages/{message_id}/attachments/{attachment_id}&lt;/CODE&gt;. That only works once the message status is &lt;CODE&gt;COMPLETED&lt;/CODE&gt;, so the &lt;CODE&gt;_and_wait&lt;/CODE&gt; variants you're using are the right fit.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Genie SDK reference: &lt;A href="https://databricks-sdk-py.readthedocs.io/en/stable/workspace/dashboards/genie.html" target="_blank"&gt;https://databricks-sdk-py.readthedocs.io/en/stable/workspace/dashboards/genie.html&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;SDK release notes (v0.120.0 added &lt;CODE&gt;enable_visualization&lt;/CODE&gt;&lt;span class="lia-unicode-emoji" title=":disappointed_face:"&gt;😞&lt;/span&gt; &lt;A href="https://github.com/databricks/databricks-sdk-py/releases" target="_blank"&gt;https://github.com/databricks/databricks-sdk-py/releases&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Genie Conversation API reference: &lt;A href="https://docs.databricks.com/api/genie/v1/conversation" target="_blank"&gt;https://docs.databricks.com/api/genie/v1/conversation&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks SDK for Python on PyPI: &lt;A href="https://pypi.org/project/databricks-sdk/" target="_blank"&gt;https://pypi.org/project/databricks-sdk/&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards,&lt;BR /&gt;Louis.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 19:17:34 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/genie-agent-sdk-visualization/m-p/170361#M2138</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-01T19:17:34Z</dc:date>
    </item>
    <item>
      <title>Re: Gemini Foundation Model endpoint unavailable in 14-day commercial trial</title>
      <link>https://community.databricks.com/t5/generative-ai/gemini-foundation-model-endpoint-unavailable-in-14-day/m-p/170359#M2137</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262701"&gt;@Remon&lt;/a&gt;&amp;nbsp;, I took a look at both internal and external documentation and here is what I found.&lt;/P&gt;
&lt;P&gt;This looks like a model availability or routing issue, not a missing custom endpoint. Seeing &lt;CODE&gt;databricks-gemini-3-8-flash&lt;/CODE&gt; under &lt;CODE&gt;system.ai&lt;/CODE&gt; doesn't mean it's enabled for your workspace. That catalog is a global list of everything Databricks hosts anywhere. Whether a callable model service exists in a given workspace depends on region, cross-Geo settings, serving capacity, and your permissions. The docs say this plainly on the Unity Gateway models page, so NotFound here is consistent with "listed but not materialized," not a broken workspace.&lt;/P&gt;
&lt;P&gt;The most likely cause is cross-geography routing. Gemini 3.8 Flash is hosted on a global endpoint, and its model page states it requires cross geography routing to be enabled. Your Tokyo workspace is in the Asia Geo, so requests have to be allowed to leave that Geo. It's normally on by default outside the US and EU, but it's off if a compliance security profile is enabled, and it's a per-workspace toggle that's easy to miss. To check:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Sign in to the account console (&lt;A href="https://accounts.cloud.databricks.com/" target="_blank"&gt;https://accounts.cloud.databricks.com/&lt;/A&gt;) as an account admin. If you set up the trial, that should be you.&lt;/LI&gt;
&lt;LI&gt;Click Workspaces and open your Tokyo workspace.&lt;/LI&gt;
&lt;LI&gt;Open the Security and compliance tab.&lt;/LI&gt;
&lt;LI&gt;Make sure "Enforce data processing within workspace Geography for Designated Services" is turned off.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Then look for the &lt;CODE&gt;system.ai.gemini-3-8-flash&lt;/CODE&gt; model service in Unity Gateway and confirm your user can query it.&lt;/P&gt;
&lt;P&gt;Second, check the request shape. There are three API paths now, and mixing identifiers across them produces exactly the error you quoted:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Unity Gateway unified API: POST to &lt;CODE&gt;https://&amp;lt;workspace-url&amp;gt;/ai-gateway/mlflow/v1/chat/completions&lt;/CODE&gt; with &lt;CODE&gt;"model": "system.ai.gemini-3-8-flash"&lt;/CODE&gt;&lt;/LI&gt;
&lt;LI&gt;Native Gemini API: &lt;CODE&gt;https://&amp;lt;workspace-url&amp;gt;/ai-gateway/gemini/v1beta/models/system.ai.gemini-3-8-flash:generateContent&lt;/CODE&gt;&lt;/LI&gt;
&lt;LI&gt;Legacy Model Serving: &lt;CODE&gt;https://&amp;lt;workspace-url&amp;gt;/serving-endpoints/databricks-gemini-3-8-flash/invocations&lt;/CODE&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Sending &lt;CODE&gt;system.ai.gemini-3-8-flash&lt;/CODE&gt; to &lt;CODE&gt;/serving-endpoints/&lt;/CODE&gt;, or &lt;CODE&gt;databricks-gemini-3-8-flash&lt;/CODE&gt; to &lt;CODE&gt;/ai-gateway/&lt;/CODE&gt;, fails with NotFound even when everything is provisioned correctly. Pay-per-token models are managed services, so you shouldn't need to create any endpoint yourself.&lt;/P&gt;
&lt;P&gt;A quick isolation test: try &lt;CODE&gt;system.ai.gemini-3-7-flash&lt;/CODE&gt; or &lt;CODE&gt;system.ai.claude-sonnet-4-6&lt;/CODE&gt; first. Both are listed for &lt;CODE&gt;ap-northeast-1&lt;/CODE&gt; without the cross-Geo requirement. If those work and 3.8 Flash doesn't, it's the Geo setting. If none work, it's the request shape or entitlements.&lt;/P&gt;
&lt;P&gt;Your four questions, in order:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Trial availability: neither the public docs nor our internal references confirm or deny a trial-specific restriction on this model. If region, routing, identifier, and permissions all check out and it's still absent, that's the point where Databricks would need to verify trial entitlement or backend provisioning.&lt;/LI&gt;
&lt;LI&gt;Listed but not callable: global catalog listing versus regional availability, as above.&lt;/LI&gt;
&lt;LI&gt;Entitlements or provisioning: you need Unity Catalog enabled, the Workspace access entitlement on your user, and the cross-Geo setting. No ticket should be needed.&lt;/LI&gt;
&lt;LI&gt;Workspace setting: the account console toggle above. A metastore admin can also restrict which &lt;CODE&gt;system.ai&lt;/CODE&gt; models users may query, though on a fresh trial that's unlikely to have changed.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;One caveat. The region availability table lists &lt;CODE&gt;databricks-gemini-3-8-flash&lt;/CODE&gt; for &lt;CODE&gt;ap-northeast-1&lt;/CODE&gt; without the cross-Geo marker, while the model's own page says it requires it. The two pages aren't perfectly consistent, so the toggle is the practical test. If you flip it, fix the request shape, and still get NotFound, reply here with your cloud and region, the exact identifier and API path, a timestamp, and the sanitized error (no tokens or credentials, please). I'll take another look.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Available models in Unity Gateway (note on &lt;CODE&gt;system.ai&lt;/CODE&gt; listing): &lt;A href="https://docs.databricks.com/aws/en/machine-learning/model-serving/foundation-model-overview" target="_blank"&gt;https://docs.databricks.com/aws/en/machine-learning/model-serving/foundation-model-overview&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Gemini 3.8 Flash model page: &lt;A href="https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/supported-models#gemini-3-8-flash" target="_blank"&gt;https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/supported-models#gemini-3-8-flash&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Model availability by region: &lt;A href="https://docs.databricks.com/aws/en/unity-gateway/model-region-availability" target="_blank"&gt;https://docs.databricks.com/aws/en/unity-gateway/model-region-availability&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks Geos and cross-Geo processing: &lt;A href="https://docs.databricks.com/aws/en/resources/databricks-geos#cross-geo-processing" target="_blank"&gt;https://docs.databricks.com/aws/en/resources/databricks-geos#cross-geo-processing&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Query model services through Unity Gateway: &lt;A href="https://docs.databricks.com/aws/en/ai-gateway/query-model-services" target="_blank"&gt;https://docs.databricks.com/aws/en/ai-gateway/query-model-services&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Query with the Google Gemini API: &lt;A href="https://docs.databricks.com/aws/en/machine-learning/model-serving/query-gemini-api" target="_blank"&gt;https://docs.databricks.com/aws/en/machine-learning/model-serving/query-gemini-api&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Use foundation models (legacy serving-endpoint query options): &lt;A href="https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-models" target="_blank"&gt;https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-models&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Regards,&lt;BR /&gt;Louis.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 19:06:50 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/gemini-foundation-model-endpoint-unavailable-in-14-day/m-p/170359#M2137</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-10-01T19:06:50Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170346#M2136</link>
      <description>&lt;P&gt;Hello&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/262617"&gt;@ErikVetter&lt;/a&gt;.&lt;/P&gt;&lt;P&gt;Thank you so much; that worked like magic.&lt;/P&gt;&lt;P&gt;Regards&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 16:08:42 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170346#M2136</guid>
      <dc:creator>HOguns</dc:creator>
      <dc:date>2026-10-01T16:08:42Z</dc:date>
    </item>
    <item>
      <title>Gemini Foundation Model endpoint unavailable in 14-day commercial trial</title>
      <link>https://community.databricks.com/t5/generative-ai/gemini-foundation-model-endpoint-unavailable-in-14-day/m-p/170335#M2135</link>
      <description>&lt;P&gt;Hi,&lt;/P&gt;&lt;P&gt;I am using a 14-day Databricks commercial trial with promotional credits on an AWS Tokyo workspace&lt;/P&gt;&lt;P&gt;I am specifically trying to access the documented "databricks-gemini-3-8-flash" Foundation Model, but it is not available through Unity Gateway / Model Serving in my workspace.&lt;/P&gt;&lt;P&gt;What I have already checked:&lt;/P&gt;&lt;P&gt;- "databricks-gemini-3-8-flash" appears under the "system.ai" catalog / registered models.&lt;/P&gt;&lt;P&gt;- My workspace is in the AWS Tokyo region.&lt;/P&gt;&lt;P&gt;- I checked the relevant workspace AI and feature settings.&lt;/P&gt;&lt;P&gt;- I checked Unity Gateway / Model Serving, but I cannot find a Gemini model service or callable endpoint.&lt;/P&gt;&lt;P&gt;- I also tried calling the model endpoint directly through the API.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The API returns:&lt;/P&gt;&lt;P&gt;"NotFound: The given endpoint does not exist, please retry after checking the specified model and version deployment exists."&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;I contacted Databricks Support, but they informed me that my account is classified as an individual account without a support contract and recommended that I post the issue here in the Community.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Could someone from Databricks please clarify:&lt;/P&gt;&lt;P&gt;1. Is "databricks-gemini-3-8-flash" available on the 14-day commercial trial?&lt;/P&gt;&lt;P&gt;2. If it is available, why does the model appear in "system.ai" but not as a callable endpoint in my workspace?&lt;/P&gt;&lt;P&gt;3. Does my workspace require any additional entitlement, backend provisioning, or activation?&lt;/P&gt;&lt;P&gt;4. Is there any workspace setting I need to enable, or does Databricks need to provision the model for my account?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Thank you.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 13:48:08 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/gemini-foundation-model-endpoint-unavailable-in-14-day/m-p/170335#M2135</guid>
      <dc:creator>Remon</dc:creator>
      <dc:date>2026-10-01T13:48:08Z</dc:date>
    </item>
    <item>
      <title>Routing between GLM 5.3 Flash and GLM 5.3 with ai_decide: a cheap router in front of ai_query</title>
      <link>https://community.databricks.com/t5/generative-ai/routing-between-glm-5-3-flash-and-glm-5-3-with-ai-decide-a-cheap/m-p/170333#M2134</link>
      <description>&lt;P&gt;I hope this helps for those who are trying to keep LLM costs under control on batch workloads. I have been looking at ai_decide (Beta) as a router rather than as a classifier, and it fits that job well. Sharing the pattern and the open questions, since I have not benchmarked it at scale yet.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;THE IDEA&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Most rows in a typical batch job (ticket triage, review summaries, document extraction) do not need the big model. A small, fast decision step can look at each row and choose which model should handle it. ai_decide takes text plus a set of questions and returns a choice, a probability or a score, and it is documented as faster and cheaper than a general LLM for decisions. The Databricks announcement lists "routing prompts to the right model" as a use case.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Endpoints (from the supported models page; re-check in your workspace with WorkspaceClient().serving_endpoints.list(), because names change quickly):&lt;/P&gt;&lt;P&gt;- databricks-glm-5-3-flash : cheaper, multimodal, reasoning always on&lt;/P&gt;&lt;P&gt;- databricks-glm-5-3 : the full model, text only, reasoning effort configurable&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;STEP 1: ask ai_decide for a tier and a confidence&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;ai_decide(state, questions, options) returns a VARIANT. For a choice question you get the chosen label, per-label probabilities and a confidence.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;SELECT id, prompt,&lt;/P&gt;&lt;P&gt;&amp;nbsp; ai_decide(&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; prompt,&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; '{&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; "tier": {&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; "type": "choice",&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; "instructions": "Which model tier is needed to answer this request well?",&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; "criteria": {&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; "flash": "Short factual lookup, simple extraction, classification, or rewriting. No multi-step reasoning.",&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; "full": "Multi-step reasoning, ambiguous or conflicting inputs, long synthesis, code generation, or anything where an error is costly."&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp; }&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; }&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; }',&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; map('version', '1.0')&lt;/P&gt;&lt;P&gt;&amp;nbsp; ) AS d&lt;/P&gt;&lt;P&gt;FROM prompts_to_process;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Pull the fields out of the VARIANT:&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;d:response:answers:tier:choice::string -- 'flash' or 'full'&lt;/P&gt;&lt;P&gt;d:response:answers:tier:confidence::double -- how sure the router is&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;STEP 2: send each row to the model it was routed to&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;ai_query needs the endpoint name as a constant, so you cannot pass the routed name in as a column value. Two options: (a) a CASE with one ai_query call per branch, or (b) two statements, one filtered on tier = 'flash' and one on tier = 'full', writing into the same results table. I prefer (b): each batch is homogeneous, you can size and monitor each endpoint separately, and you do not depend on how the engine evaluates both branches of a CASE.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;CREATE OR REPLACE TABLE routed AS&lt;/P&gt;&lt;P&gt;SELECT id, prompt,&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;d:response:answers:tier:choice::string AS tier,&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; &amp;nbsp;d:response:answers:tier:confidence::double AS conf&lt;/P&gt;&lt;P&gt;FROM (SELECT id, prompt, ai_decide(prompt, '&amp;lt;questions json from step 1&amp;gt;', map('version','1.0')) AS d&lt;/P&gt;&lt;P&gt;&amp;nbsp; &amp;nbsp; &amp;nbsp; FROM prompts_to_process);&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;-- cheap path&lt;/P&gt;&lt;P&gt;INSERT INTO results&lt;/P&gt;&lt;P&gt;SELECT id, 'flash' AS model, ai_query('databricks-glm-5-3-flash', prompt) AS answer&lt;/P&gt;&lt;P&gt;FROM routed WHERE tier = 'flash' AND conf &amp;gt;= 0.7;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;-- everything else goes to the full model, including low-confidence rows&lt;/P&gt;&lt;P&gt;INSERT INTO results&lt;/P&gt;&lt;P&gt;SELECT id, 'full' AS model, ai_query('databricks-glm-5-3', prompt) AS answer&lt;/P&gt;&lt;P&gt;FROM routed WHERE tier = 'full' OR conf &amp;lt; 0.7;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;The design choice I think matters most: send low-confidence rows to the full model. A router that is unsure should fail expensive, not fail wrong. Tune the 0.7 threshold on your own data.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;THINGS TO CHECK BEFORE TRUSTING THIS IN PRODUCTION&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;1. Label a sample (a few hundred rows) by running both models and comparing. The router only saves money if flash answers are good enough on the rows it keeps. Measure answer quality per tier, not just router accuracy.&lt;/P&gt;&lt;P&gt;2. Watch the cost of the router itself. ai_decide is cheap per call, but on very short prompts it may be a meaningful share of the flash call.&lt;/P&gt;&lt;P&gt;3. Log tier, confidence and model per row so you can see the real flash/full split and recompute savings.&lt;/P&gt;&lt;P&gt;4. ai_decide is Beta, only available in some regions, and not on Databricks SQL Classic, so check your warehouse type and region first.&lt;/P&gt;&lt;P&gt;5. Re-verify endpoint names and pricing. Both are pay-per-token endpoints and the lineup moves fast.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;WHAT I HAVE NOT TESTED&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;I have not benchmarked routing accuracy or the real cost split at scale, so treat the threshold and the criteria wording as a starting point. The criteria text moves results the most, so it is worth iterating on it with a labelled sample.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;QUESTIONS FOR THE COMMUNITY&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;- Has anyone tried ai_decide as a router at volume? What flash/full split did you land on?&lt;/P&gt;&lt;P&gt;- Any experience with score-type questions (a 0 to 2 difficulty score) versus a choice question for routing? I picked choice because it also returns a confidence I can threshold on.&lt;/P&gt;&lt;P&gt;- Are you routing on prompt content alone, or also on metadata such as customer tier or SLA?&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;References: ai_decide SQL function reference and announcement blog, and the Foundation Model APIs supported models page on docs.databricks.com.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 13:21:43 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/routing-between-glm-5-3-flash-and-glm-5-3-with-ai-decide-a-cheap/m-p/170333#M2134</guid>
      <dc:creator>DoTA</dc:creator>
      <dc:date>2026-10-01T13:21:43Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170303#M2133</link>
      <description>&lt;P&gt;Hey, I had a simillar problem, with the same Error Code.&lt;BR /&gt;I asked Codex to fix it and the solution is very simple.&lt;/P&gt;&lt;P&gt;In the Genie Code window, next to the send button, change from auto to low.&lt;BR /&gt;I dont know, this seems like a bug since some weeks to me.&lt;/P&gt;</description>
      <pubDate>Thu, 01 Oct 2026 08:41:34 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170303#M2133</guid>
      <dc:creator>ErikVetter</dc:creator>
      <dc:date>2026-10-01T08:41:34Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170275#M2132</link>
      <description>&lt;P&gt;Alright, thank you. I will have to delete the workspace now because its running on a pay as you go and I am still stuck.&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 17:39:27 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170275#M2132</guid>
      <dc:creator>HOguns</dc:creator>
      <dc:date>2026-09-30T17:39:27Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170243#M2131</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/261901"&gt;@HOguns&lt;/a&gt;,&lt;BR /&gt;&lt;BR /&gt;I am not an engineer, I cover the delivery side of Genie.&amp;nbsp;&lt;BR /&gt;&lt;BR /&gt;However, I passed it to one of our internal teams. Let's see if they pick it up. I am not sure how engineers pick it up from here&amp;nbsp;&lt;BR /&gt;&lt;BR /&gt;&lt;SPAN&gt;&amp;nbsp;In the meantime, can you also please send it to&amp;nbsp;&lt;/SPAN&gt;&lt;A href="mailto:databricks-community@databricks.com" target="_blank"&gt;databricks-community@databricks.com&lt;/A&gt;&lt;SPAN&gt;, so it is not missed out on&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Wed, 30 Sep 2026 12:17:36 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170243#M2131</guid>
      <dc:creator>Valeria_Koz_DBX</dc:creator>
      <dc:date>2026-09-30T12:17:36Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170184#M2130</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/221494"&gt;@Valeria_Koz_DBX&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;I contacted the Help Center via email since I am unable to do point 1 above to open a case. The response from your team via email was to post it here, and one of the engineers would help instead.&lt;/P&gt;&lt;P&gt;Regards&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 18:01:56 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170184#M2130</guid>
      <dc:creator>HOguns</dc:creator>
      <dc:date>2026-09-29T18:01:56Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170177#M2129</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/261901"&gt;@HOguns&lt;/a&gt;&amp;nbsp;&lt;BR /&gt;&lt;BR /&gt;Did you posted here, or Contacted Help Center with all of the details above?&lt;BR /&gt;&lt;BR /&gt;Ideally, you want to follow the below steps&lt;/P&gt;
&lt;UL&gt;
&lt;LI class="p8i6j0c"&gt;&lt;STRONG&gt;Use the Databricks Help Center&lt;/STRONG&gt; if you can access it, and open a case for &lt;STRONG&gt;Azure Databricks → Genie Code/Assistant&lt;/STRONG&gt;.&lt;/LI&gt;
&lt;LI class="p8i6j0c"&gt;If Help Center access is denied, email &lt;STRONG&gt;help@databricks.com&lt;/STRONG&gt; from the same address used for the Databricks account. Help Center access may require an authorized support contact or support entitlement.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 16:44:25 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170177#M2129</guid>
      <dc:creator>Valeria_Koz_DBX</dc:creator>
      <dc:date>2026-09-29T16:44:25Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170175#M2128</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/221494"&gt;@Valeria_Koz_DBX&lt;/a&gt;,&lt;/P&gt;&lt;P&gt;Thank you for your prompt feedback; I have done exactly as you highlighted in No. 1 with the service principal name "&lt;SPAN&gt;mondayhenry31_gmail.com#EXT#@mondayhenry31gmail.onmicrosoft.com" and I have used this to sign in using the in-Private/incognito mode, but I still hit the same impasse. Based on No 2.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;Databricks account ID - 06348851-c57a-4116-9779-6baae571c2a1&lt;/P&gt;&lt;P&gt;Workspace ID and workspace URL -&amp;nbsp;7405617987563324&lt;/P&gt;&lt;P&gt;Azure region - East US&lt;/P&gt;&lt;P&gt;The exact identity/UPN used to sign in -&amp;nbsp;&lt;SPAN&gt;mondayhenry31_gmail.com#EXT#@mondayhenry31gmail.onmicrosoft.com&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;Approximate time of the failure and a screenshot of the error - 11:40 Canada Time EDT&lt;/P&gt;&lt;P&gt;Confirmation that the user identity and workspace assignment were checked - Yes it was checked&lt;/P&gt;&lt;P&gt;Below I have provided the workspace&amp;nbsp; details as well&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Name&lt;/SPAN&gt;&lt;/P&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;Status&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;Resource group&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;Region&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;Subscription&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;Created&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;Metastore&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&amp;nbsp;&lt;/DIV&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;&lt;A target="_blank"&gt;adb-dp750&lt;/A&gt;&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;&lt;SPAN class=""&gt;Running&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;rg-dp750&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;eastus&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;06348851-c57a-4116-9779-6baae571c2a1&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;yesterday at 3:37 PM&lt;/SPAN&gt;&lt;/DIV&gt;&lt;DIV class=""&gt;&lt;SPAN class=""&gt;&lt;A target="_blank"&gt;metastore_azure_eastus&lt;/A&gt;&lt;/SPAN&gt;&lt;SPAN class=""&gt;&lt;A class="" href="https://adb-7405617987563324.4.azuredatabricks.net/aad/auth" target="_blank" rel="noopener noreferrer"&gt;Open&lt;/A&gt;&lt;/SPAN&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;/DIV&gt;&lt;P&gt;Regard&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 15:53:35 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170175#M2128</guid>
      <dc:creator>HOguns</dc:creator>
      <dc:date>2026-09-29T15:53:35Z</dc:date>
    </item>
    <item>
      <title>Re: WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170170#M2127</link>
      <description>&lt;P&gt;Hello &lt;A href="https://community.databricks.com/t5/user/viewprofilepage/user-id/261901" target="_blank"&gt;@HOguns&lt;/A&gt;&amp;nbsp;,&lt;BR /&gt;&lt;BR /&gt;Long response, but it is worth it.&lt;BR /&gt;&lt;BR /&gt;Thank you for sharing the screenshots. There appear to be two separate issues: the Genie Code daily-token allowance and Help Center access.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;1. Confirm the Databricks identity being used&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The screenshots show several external identities for the same person, including mondayhenry31@gmail.com, an Azure guest identity containing #ext#, and onmicrosoft.com identities. The &lt;STRONG&gt;Account admin&lt;/STRONG&gt; role appears to be assigned to the #ext# identity, not clearly to the plain Gmail identity.&lt;/P&gt;
&lt;P&gt;Please ask your Microsoft Entra ID administrator to confirm which user principal name (UPN) is used when you sign in to Azure Databricks. Then:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Sign out of Databricks and Microsoft accounts in the browser, or use an incognito window.&lt;/LI&gt;
&lt;LI&gt;Sign in using the same identity that has the Databricks &lt;STRONG&gt;Account admin&lt;/STRONG&gt; role.&lt;/LI&gt;
&lt;LI&gt;Confirm that this identity is assigned to the affected workspace.&lt;/LI&gt;
&lt;LI&gt;If workspace administration is required, confirm that it also has the &lt;STRONG&gt;Workspace admin&lt;/STRONG&gt;&lt;SPAN&gt; role.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Account-admin access and workspace access are separate permissions. Azure Databricks also recommends avoiding duplicate identities and using one identity-provisioning method. Do not delete the duplicate entries until the canonical Entra ID identity has been confirmed.&lt;/P&gt;
&lt;P&gt;References: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/admin/admin-concepts" rel="noopener noreferrer nofollow" target="_blank"&gt;Azure Databricks account administration&lt;/A&gt;, &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/admin/users-groups/users" rel="noopener noreferrer nofollow" target="_blank"&gt;Manage users in Azure Databricks&lt;/A&gt;, and &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/admin/users-groups/automatic-identity-management/" rel="noopener noreferrer nofollow" target="_blank"&gt;Automatic identity management&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;2. The budget does not create the daily token allowance&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The screenshot confirms that Genie Code is displaying the error:&lt;/P&gt;
&lt;P class=""&gt;This workspace has no daily token allowance for the assistant. Ask your workspace admin to enable it, then try again.&lt;/P&gt;
&lt;P&gt;A Genie budget is a usage-cost control; creating a budget does not by itself provision or enable the assistant’s daily token allowance. If a budget is configured, verify that it is scoped to the Genie product using the documented Unity Gateway resource and databricks-product = genie tag.&lt;/P&gt;
&lt;P&gt;After correcting the sign-in identity and workspace assignment, start a new Genie Code session and test again. If the same error remains, this requires a Databricks-side account/workspace entitlement or provisioning check rather than another change in User Management.&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;When contacting Databricks, include:&lt;/LI&gt;
&lt;LI&gt;Databricks account ID&lt;/LI&gt;
&lt;LI&gt;Workspace ID and workspace URL&lt;/LI&gt;
&lt;LI&gt;Azure region&lt;/LI&gt;
&lt;LI&gt;The exact identity/UPN used to sign in&lt;/LI&gt;
&lt;LI&gt;Approximate time of the failure and a screenshot of the error&lt;/LI&gt;
&lt;LI&gt;Confirmation that the user identity and workspace assignment were checked&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Reference: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/genie-code/" rel="noopener noreferrer nofollow" target="_blank"&gt;Genie Code on Azure Databricks&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;3. Help Center access is separate from workspace administration&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;The message stating that the user does not have access to Help Center is not necessarily caused by missing workspace permissions. Help Center access depends on the organization’s Databricks support entitlement and whether the exact email address is registered as an authorized support contact.&lt;/P&gt;
&lt;P&gt;The customer should:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Sign in to Help Center with the exact email associated with the Databricks support entitlement.&lt;/LI&gt;
&lt;LI&gt;If mondayhenry31@gmail.com is not the authorized support contact, ask the account owner or support administrator to add the correct email address.&lt;/LI&gt;
&lt;LI&gt;If this is a course, trial, or learning environment without a support contract, use the course support channel or Databricks Community; Help Center case access may not be included.&lt;/LI&gt;
&lt;LI&gt;If access is still denied, email help@databricks.com with the account ID, workspace ID, workspace URL, Azure region, login identity, and the Help Center error.&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;Reference: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/resources/support" rel="noopener noreferrer nofollow" target="_blank"&gt;Azure Databricks support&lt;/A&gt;.&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Summary&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;A separate Azure subscription is not normally required just to use the workspace or assign Databricks administrator roles. However, the screenshots suggest that the account has an identity/role mismatch that should be corrected first. If the Genie Code error continues after using the canonical identity, Databricks Support must verify the workspace’s Genie Code daily allowance and provisioning. Help Center access is a separate support entitlement from workspace administration.&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 15:04:31 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170170#M2127</guid>
      <dc:creator>Valeria_Koz_DBX</dc:creator>
      <dc:date>2026-09-29T15:04:31Z</dc:date>
    </item>
    <item>
      <title>WORKSPACE HAS NO DAILY TOKEN ALLOWANCE</title>
      <link>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170156#M2125</link>
      <description>&lt;P&gt;Hello Team,&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;When I try to debug and troubleshoot this error: "This workspace has no daily token allowance for the assistant. Ask your workspace admin to enable it, then try again." I am the Admin, but I still don't see any way to resolve this issue, I have tried creating a budget with the product genie, I have also navigated to the User Management and on there as shown in on of the snapshot the users and the Admin still I keep hitting that error prompt when I use it alongside a notebook in the workspace.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;After chatting with the agent on your website, when I try to access the Help Center, I keep getting this error prompt: "You currently do not have access to Help Center. Please reach out to your admin or send an email to help@databricks.com" for my email address: "&lt;A href="mailto:mondayhenry31@gmail.com" target="_blank" rel="noopener"&gt;mondayhenry31@gmail.com&lt;/A&gt;".&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Is it that I need a separate subscription from my Azure on Databricks and like a separate Admin access which I can't seem to understand. Please help, as I am trying to go through a hands-on course project, and I am stuck here.&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Regards&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 13:08:28 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/workspace-has-no-daily-token-allowance/m-p/170156#M2125</guid>
      <dc:creator>HOguns</dc:creator>
      <dc:date>2026-09-29T13:08:28Z</dc:date>
    </item>
    <item>
      <title>Re: Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused</title>
      <link>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170150#M2124</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/139171"&gt;@qduan&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Also do check the backend state before re-creating the assistant. I had tested a similar failure path and saw the assistant remain in CREATING while the file source stayed in UPDATING, with no error_info surfaced yet.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Few useful checks:&lt;/STRONG&gt;&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;import requests
req = requests.get(
    f"https://{workspace_url}/api/2.1/{assistant_name}",
    headers=headers
)
print(req.json())
req = requests.get(
    f"https://{workspace_url}/api/2.1/{assistant_name}/knowledge-sources",
    headers=headers
)
print(req.json())&lt;/LI-CODE&gt;&lt;P&gt;Look for:&lt;BR /&gt;- &lt;STRONG&gt;assistant state&lt;/STRONG&gt; (CREATING, ACTIVE, FAILED)&lt;BR /&gt;- &lt;STRONG&gt;error_info&lt;/STRONG&gt;&lt;BR /&gt;- &lt;STRONG&gt;source state&lt;/STRONG&gt; (UPDATING, UPDATED, FAILED_UPDATE)&lt;BR /&gt;- &lt;STRONG&gt;knowledge_cutoff_time&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Also check the Audit Log:&lt;/STRONG&gt;&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;SELECT
  event_time,
  action_name,
  response.status_code,
  response.error_message,
  request_params
FROM system.access.audit
WHERE service_name = 'knowledgeAssistant'
  AND event_time &amp;gt;= current_timestamp() - INTERVAL 1 HOUR
ORDER BY event_time DESC&lt;/LI-CODE&gt;&lt;P&gt;In my case when I tested, the source stayed stuck in UPDATING and the assistant in CREATING even though the API GET calls returned 200 and no later sync error appeared in the audit log.&amp;nbsp;&amp;nbsp;&lt;/P&gt;&lt;P&gt;So if the UI says the agent cannot load, checking the assistant/source states through the API is useful because the backend may be stuck in provisioning without surfacing a clear failure event yet.&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 11:32:33 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170150#M2124</guid>
      <dc:creator>data_pulse</dc:creator>
      <dc:date>2026-09-29T11:32:33Z</dc:date>
    </item>
    <item>
      <title>Re: Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused</title>
      <link>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170132#M2123</link>
      <description>&lt;P&gt;Interesting issue—especially the index compatibility part. Hopefully there’s a supported way to recover the agent without re-indexing 2,000+ files.&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 09:35:26 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170132#M2123</guid>
      <dc:creator>ThiamLee</dc:creator>
      <dc:date>2026-09-29T09:35:26Z</dc:date>
    </item>
    <item>
      <title>Re: Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused</title>
      <link>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170129#M2122</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/139171"&gt;@qduan&lt;/a&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Genie is right regarding the embedding models. When Knowledge Assistant ingests files directly, it creates a managed, internal index using databricks-agent-bricks-embedding-v1. Because that model isn't on the supported list for external AI Search indexes (databricks-gte-large-en, databricks-bge-large-en, or databricks-qwen3-embedding-0-6b), it won't appear as an available source for a new assistant as its a design restriction. There is not a supported way to use that auto generated internal index and repurpose it for new assistants. You may prepare an index and embed the files using one of those three supported models if you plan to use index as an source.&lt;/P&gt;&lt;P&gt;For the load agent error, its likely a platform side issue on the agent's serving endpoint and not due to misconfigurations at your end. You can check - Serving in the platform and verify if the agent's endpoint exists and its status currently. If inference table is enabled, you can run a quick query filtering on status_code to get the exact error details if available. You can inspect the raw metadata using the SDK to see the state the backend agent&lt;/P&gt;&lt;P&gt;Given the volume of 2000+ mixed files and the likelihood of a service side issue during creation, you can reach Databricks support. Share them workspace URL, the creation time and the agent name and they can check it.&lt;/P&gt;&lt;P&gt;Recreating the assistant and reingesting the files is the path forward since that internal index cannot be used.&lt;/P&gt;</description>
      <pubDate>Tue, 29 Sep 2026 09:19:10 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170129#M2122</guid>
      <dc:creator>balajij8</dc:creator>
      <dc:date>2026-09-29T09:19:10Z</dc:date>
    </item>
    <item>
      <title>Re: Salesforce Hosted MCP + UC HTTP connection — stops working after ~1 hour?</title>
      <link>https://community.databricks.com/t5/generative-ai/salesforce-hosted-mcp-uc-http-connection-stops-working-after-1/m-p/170102#M2121</link>
      <description>&lt;UL&gt;&lt;LI&gt;&lt;STRONG&gt;Exact error&lt;/STRONG&gt;: HTTP 401, JSON-RPC -32006, "You're not signed in"&lt;/LI&gt;&lt;/UL&gt;</description>
      <pubDate>Tue, 29 Sep 2026 01:04:58 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/salesforce-hosted-mcp-uc-http-connection-stops-working-after-1/m-p/170102#M2121</guid>
      <dc:creator>iankizer</dc:creator>
      <dc:date>2026-09-29T01:04:58Z</dc:date>
    </item>
    <item>
      <title>Salesforce Hosted MCP + UC HTTP connection — stops working after ~1 hour?</title>
      <link>https://community.databricks.com/t5/generative-ai/salesforce-hosted-mcp-uc-http-connection-stops-working-after-1/m-p/170094#M2120</link>
      <description>&lt;P class=""&gt;We're wiring&amp;nbsp;&lt;SPAN&gt;Salesforce Hosted MCP&lt;/SPAN&gt;&amp;nbsp;(read-only SObject server on&amp;nbsp;&lt;SPAN class=""&gt;api.salesforce.com&lt;/SPAN&gt;) through a&amp;nbsp;&lt;SPAN&gt;Unity Catalog HTTP connection&lt;/SPAN&gt;&amp;nbsp;and an&amp;nbsp;&lt;SPAN&gt;MCP service&lt;/SPAN&gt;, using&amp;nbsp;&lt;SPAN&gt;OAuth User-to-Machine Per User&lt;/SPAN&gt;.&lt;/P&gt;&lt;P class=""&gt;Login on the MCP service works and tools respond at first. After about&amp;nbsp;&lt;SPAN&gt;an hour&lt;/SPAN&gt;, MCP calls fail until we log in again on the service.&lt;/P&gt;&lt;P class=""&gt;On the HTTP connection,&amp;nbsp;&lt;SPAN&gt;access token expiration&lt;/SPAN&gt;&amp;nbsp;shows&amp;nbsp;&lt;SPAN&gt;“Not provided by provider.”&lt;/SPAN&gt;&amp;nbsp;Other OAuth HTTP connections in the same workspace (different vendors) show an expiration and don’t hit this hourly cliff.&lt;/P&gt;&lt;P class=""&gt;Salesforce side is an&amp;nbsp;&lt;SPAN&gt;External Client App&lt;/SPAN&gt;&amp;nbsp;with&amp;nbsp;&lt;SPAN&gt;mcp_api&lt;/SPAN&gt;&amp;nbsp;/&amp;nbsp;&lt;SPAN&gt;refresh_token&lt;/SPAN&gt;, PKCE, and&amp;nbsp;&lt;SPAN&gt;JWT-based access tokens&lt;/SPAN&gt;&amp;nbsp;(as in Salesforce’s Hosted MCP docs). Callback uses the standard Databricks&amp;nbsp;&lt;SPAN&gt;&lt;SPAN class=""&gt;/login/oauth/&lt;/SPAN&gt;&lt;SPAN class=""&gt;http.html&lt;/SPAN&gt;&lt;/SPAN&gt;&amp;nbsp;redirect.&lt;/P&gt;&lt;P class=""&gt;&lt;SPAN&gt;Has anyone got this combo stable past ~1 hour?&lt;/SPAN&gt;&lt;BR /&gt;&amp;nbsp;Is missing&amp;nbsp;&lt;SPAN&gt;expires_in&lt;/SPAN&gt;&amp;nbsp;on Salesforce token responses a known issue for UC / AI Gateway refresh? Any connection settings that actually help (token exchange method, client secret on refresh, etc.)—without turning off JWT on the ECA?&lt;/P&gt;&lt;P class=""&gt;Has anyone seen this sort of pattern and what fixed it (or if you had to escalate to support)? Or, if you've configured a Salesforce MCP in UC AI Gateway, how do have your External Client App set up so this refresh works properply&lt;/P&gt;&lt;P class=""&gt;Thanks.&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 19:30:40 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/salesforce-hosted-mcp-uc-http-connection-stops-working-after-1/m-p/170094#M2120</guid>
      <dc:creator>iankizer</dc:creator>
      <dc:date>2026-09-28T19:30:40Z</dc:date>
    </item>
    <item>
      <title>Re: Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused</title>
      <link>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170082#M2119</link>
      <description>&lt;P&gt;Hello &lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/139171"&gt;@qduan&lt;/a&gt;&amp;nbsp;, I did some digging and here is what I found.&lt;/P&gt;
&lt;P&gt;Genie has the embedding piece right. The Knowledge Assistant docs list exactly three supported models for a bring-your-own AI Search index: databricks-gte-large-en, databricks-bge-large-en, and databricks-qwen3-embedding-0-6b. An index built automatically from a files source uses a Databricks-managed default model that isn't on that list, so it won't show up in the picker for a new assistant. That's documented behavior, not a bug on your end. I couldn't find anything public that promises reuse of that generated index without re-embedding, so I'd treat it as unsupported today.&lt;/P&gt;
&lt;P&gt;That restriction doesn't explain the "We could not load the agent" error, though. Those are two separate problems, and the load failure needs actual evidence. On logs and diagnosis:&lt;/P&gt;
&lt;OL&gt;
&lt;LI&gt;Query the audit log system table for service_name = knowledgeAssistant. Assistant and knowledge-source create, update, and sync events are logged there, so a failed or stuck sync should be visible.&lt;/LI&gt;
&lt;LI&gt;Call the Knowledge Assistants API (SDK or REST) to get the agent and list its sources. If the API returns a status, you've bypassed whatever the UI is choking on.&lt;/LI&gt;
&lt;LI&gt;In Catalog Explorer, check the generated index and its AI Search endpoint. Under Serving, check the agent's endpoint. Either one offline or failed will break the agent page.&lt;/LI&gt;
&lt;LI&gt;Confirm the original volume still exists and your user still has access to it.&lt;/LI&gt;
&lt;LI&gt;Open browser dev tools, reload the agent page, and grab the failing request and response from the Network tab. That payload is far more specific than the UI message.&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;Then open a Databricks Support case with the assistant ID, workspace ID, region, failure time, and whatever the audit log and network call gave you. A generic load error alone isn't enough to tell a permissions problem from an ingestion or service-side one. Don't delete the assistant or its source while this is open. Deleting an assistant removes everything associated with it from default storage.&lt;/P&gt;
&lt;P&gt;If support can't revive it, you're looking at re-indexing, either a new assistant from the same files or your own AI Search index using one of the three supported models. The upside of building your own is that multiple assistants can share it and it updates automatically with no manual sync. One thing to check with 2000+ mixed files: anything over 100 MB, or over 500 pages for PDF, DOC, and PPT (each slide counts as a page), is skipped during ingestion. If the original build tripped on something in that pile, it could be part of the story.&lt;/P&gt;
&lt;P&gt;References:&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;Knowledge Assistant: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/agents/agent-bricks/knowledge-assistant" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/agents/agent-bricks/knowledge-assistant&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Create AI Search endpoints and indexes: &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/ai-search/create-ai-search" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/ai-search/create-ai-search&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Audit log reference (knowledgeAssistant events): &lt;A href="https://learn.microsoft.com/en-us/azure/databricks/admin/account-settings/audit-logs" target="_blank"&gt;https://learn.microsoft.com/en-us/azure/databricks/admin/account-settings/audit-logs&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Knowledge Assistants API: &lt;A href="https://docs.databricks.com/api/workspace/knowledgeassistants" target="_blank"&gt;https://docs.databricks.com/api/workspace/knowledgeassistants&lt;/A&gt;&lt;/LI&gt;
&lt;LI&gt;Databricks Help Center: &lt;A href="https://help.databricks.com/s/" target="_blank"&gt;https://help.databricks.com/s/&lt;/A&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;
&lt;P&gt;Regards, Louis.&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 16:59:34 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170082#M2119</guid>
      <dc:creator>Louis_Frolio</dc:creator>
      <dc:date>2026-09-28T16:59:34Z</dc:date>
    </item>
    <item>
      <title>Knowledge Assistant fails to load after indexing; generated AI Search index cannot be reused</title>
      <link>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170042#M2118</link>
      <description>&lt;P&gt;Hi Databricks Community,&lt;/P&gt;&lt;P&gt;I created a Knowledge Assistant through &lt;STRONG&gt;Databricks → Agents&lt;/STRONG&gt;, using files as the knowledge source. The embedding and indexing were handled automatically during setup.&lt;/P&gt;&lt;P&gt;Since then, I have been unable to open the assistant. The UI displays:&lt;/P&gt;&lt;BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;We could not load the agent&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Please try again later, or contact support.&lt;/P&gt;&lt;/BLOCKQUOTE&gt;&lt;P&gt;&lt;STRONG&gt;Workaround attempted&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;To continue working, I tried creating another Knowledge Assistant and selecting the existing AI Search index as its knowledge source. However, the index did not appear as an available option.&lt;/P&gt;&lt;P&gt;I asked Genie to investigate. It reported that the automatically created index, [index_name], uses databricks-agent-bricks-embedding-v1, while Knowledge Assistant supports indexes using only:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;databricks-gte-large-en&lt;/LI&gt;&lt;LI&gt;databricks-bge-large-en&lt;/LI&gt;&lt;LI&gt;databricks-qwen3-embedding-0-6b&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;This is Genie's explanation; I have not independently confirmed that this restriction is the cause.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Expected behavior&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The original Knowledge Assistant should remain accessible after indexing. Alternatively, I would expect to be able to reuse its generated index in a new assistant without embedding and indexing the same files again.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Questions&lt;/STRONG&gt;&lt;/P&gt;&lt;OL&gt;&lt;LI&gt;How can I diagnose and resolve the “We could not load the agent” error? Are there logs I can check?&lt;/LI&gt;&lt;LI&gt;Is an index using databricks-agent-bricks-embedding-v1 intentionally unavailable as a source for another Knowledge Assistant?&lt;/LI&gt;&lt;LI&gt;Is there a supported way to recover the original assistant or reuse its existing index without re-embedding the source files?&lt;/LI&gt;&lt;/OL&gt;&lt;P&gt;&lt;STRONG&gt;Environment&lt;/STRONG&gt;&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Cloud provider / region: Azure/EMEA&lt;/LI&gt;&lt;LI&gt;Approximate creation time and time zone: 11-Sep, around 12am Amsterdam time&lt;/LI&gt;&lt;LI&gt;Source file types and approximate volume: more than 2000 files, mixed types, PDFs, pptx, etc.&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;Any guidance on recovery or index reuse would be appreciated.&lt;/P&gt;</description>
      <pubDate>Mon, 28 Sep 2026 12:00:37 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/knowledge-assistant-fails-to-load-after-indexing-generated-ai/m-p/170042#M2118</guid>
      <dc:creator>qduan</dc:creator>
      <dc:date>2026-09-28T12:00:37Z</dc:date>
    </item>
  </channel>
</rss>

