<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>All Generative AI posts</title>
    <link>https://community.databricks.com/t5/generative-ai/bd-p/GenAI-Insight-Hub</link>
    <description>All Generative AI posts</description>
    <pubDate>Sat, 15 Aug 2026 01:04:30 GMT</pubDate>
    <dc:creator>GenAI-Insight-Hub</dc:creator>
    <dc:date>2026-08-15T01:04:30Z</dc:date>
    <item>
      <title>Re: Databricks Knowledge Assistant</title>
      <link>https://community.databricks.com/t5/generative-ai/databricks-knowledge-assistant/m-p/165693#M2004</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/245057"&gt;@suryaprayaga&lt;/a&gt;&amp;nbsp;, I see the instructed-retriever-1 generates pair of query &amp;amp; filter, there can be filters that vector-index does not contain, what all metadata fields are persisted on vector-index , in case of llama-index we define the metadata columns to be persisted in vector-index right so how it works in case of KA. Wanted to understand the crux behind it.&lt;/P&gt;</description>
      <pubDate>Fri, 14 Aug 2026 19:26:44 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/databricks-knowledge-assistant/m-p/165693#M2004</guid>
      <dc:creator>IM_01</dc:creator>
      <dc:date>2026-08-14T19:26:44Z</dc:date>
    </item>
    <item>
      <title>Re: facing issue in llm models</title>
      <link>https://community.databricks.com/t5/generative-ai/facing-issue-in-llm-models/m-p/165660#M2003</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/247797"&gt;@shrm80&lt;/a&gt;&amp;nbsp; Following could be the reason:&lt;BR /&gt;- Verification delays, usually takes sometime ~ 24hrs&lt;BR /&gt;- Check AWS marketplace in Billing&lt;BR /&gt;- If all good, raise a support ticket&lt;BR /&gt;&lt;BR /&gt;In the meantime, you could test it with open-weights models (Llama).&lt;/P&gt;</description>
      <pubDate>Fri, 14 Aug 2026 07:53:57 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/facing-issue-in-llm-models/m-p/165660#M2003</guid>
      <dc:creator>Sumit_7</dc:creator>
      <dc:date>2026-08-14T07:53:57Z</dc:date>
    </item>
    <item>
      <title>Re: How to use AI for photo filter during registration like OwnMates?</title>
      <link>https://community.databricks.com/t5/generative-ai/how-to-use-ai-for-photo-filter-during-registration-like-ownmates/m-p/165659#M2002</link>
      <description>&lt;P&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/247811"&gt;@C_45&lt;/a&gt;&amp;nbsp; Great, add more references to the project.&lt;/P&gt;</description>
      <pubDate>Fri, 14 Aug 2026 07:47:38 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/how-to-use-ai-for-photo-filter-during-registration-like-ownmates/m-p/165659#M2002</guid>
      <dc:creator>Sumit_7</dc:creator>
      <dc:date>2026-08-14T07:47:38Z</dc:date>
    </item>
    <item>
      <title>How to use AI for photo filter during registration like OwnMates?</title>
      <link>https://community.databricks.com/t5/generative-ai/how-to-use-ai-for-photo-filter-during-registration-like-ownmates/m-p/165646#M2001</link>
      <description>&lt;P&gt;I am creating a website where user can upload the pic but it will be automatically filter if the picture created by AI.&lt;/P&gt;&lt;P&gt;Some social website are using it&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="1000095434.jpg" style="width: 720px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/30017i6BCDC5477D48621F/image-size/medium?v=v2&amp;amp;px=400" role="button" title="1000095434.jpg" alt="1000095434.jpg" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt; #ai #artificial&lt;/P&gt;</description>
      <pubDate>Thu, 13 Aug 2026 22:25:31 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/how-to-use-ai-for-photo-filter-during-registration-like-ownmates/m-p/165646#M2001</guid>
      <dc:creator>C_45</dc:creator>
      <dc:date>2026-08-13T22:25:31Z</dc:date>
    </item>
    <item>
      <title>facing issue in llm models</title>
      <link>https://community.databricks.com/t5/generative-ai/facing-issue-in-llm-models/m-p/165636#M2000</link>
      <description>&lt;P&gt;&lt;SPAN&gt;i have signed up to databricks using aws but when i try to access model i m getting&amp;nbsp;error&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN&gt;Error (403): PERMISSION_DENIED: The endpoint is temporarily disabled due to a Databricks-set rate limit of 0.&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 13 Aug 2026 18:03:22 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/facing-issue-in-llm-models/m-p/165636#M2000</guid>
      <dc:creator>shrm80</dc:creator>
      <dc:date>2026-08-13T18:03:22Z</dc:date>
    </item>
    <item>
      <title>Re: Testing ai_parse_document vs PyMuPDF for PDF extraction</title>
      <link>https://community.databricks.com/t5/generative-ai/testing-ai-parse-document-vs-pymupdf-for-pdf-extraction/m-p/165621#M1999</link>
      <description>&lt;P&gt;Great points here! I'm using a similar hybrid approach in production:&lt;/P&gt;&lt;P&gt;PyMuPDF runs on everything first, it extracts native text from digital PDFs (fast, free, zero AI cost). Any PDF that returns no text gets tagged as needs_ocr.&lt;/P&gt;&lt;P&gt;Only the needs_ocr documents go through ai_parse_document, scanned PDFs, complex layouts, etc. This keeps AI costs minimal.&lt;/P&gt;&lt;P&gt;For the consumption layer, instead of ai_query we chunk the extracted text (4K chars, 200 overlap) and expose it through a Gold view joined with business context. Attached that to a Genie space, users can now search inside 150K+ documents in natural language and it works surprisingly well.&lt;/P&gt;&lt;P&gt;ai_parse_document is powerful but expensive at scale. Using it only as a fallback for what PyMuPDF can't handle gives you the best of both worlds.&lt;/P&gt;</description>
      <pubDate>Thu, 13 Aug 2026 14:03:39 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/testing-ai-parse-document-vs-pymupdf-for-pdf-extraction/m-p/165621#M1999</guid>
      <dc:creator>LucasFazzi</dc:creator>
      <dc:date>2026-08-13T14:03:39Z</dc:date>
    </item>
    <item>
      <title>Re: Building an Agentic HR Front Door on Databricks</title>
      <link>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165225#M1998</link>
      <description>&lt;P&gt;Interesting to see the shift from AI assistants that answer questions to agents that can actually complete tasks. For enterprise adoption, I’m curious how teams are thinking about measuring business impact here — is the focus more on reducing manual HR workflows, improving employee experience, or increasing the accuracy and speed of decisions?&lt;/P&gt;</description>
      <pubDate>Mon, 10 Aug 2026 06:58:04 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165225#M1998</guid>
      <dc:creator>kartikchoudhary</dc:creator>
      <dc:date>2026-08-10T06:58:04Z</dc:date>
    </item>
    <item>
      <title>Re: Databricks Knowledge Assistant</title>
      <link>https://community.databricks.com/t5/generative-ai/databricks-knowledge-assistant/m-p/165136#M1997</link>
      <description>&lt;P&gt;Yes this is very much implemented, however I haven't tested it much as much I have tested in Grafana's Assistant. The new feature&amp;nbsp;&lt;SPAN&gt;Instructed-Retriever-1 is now conveniently called as Knowledge Assistant.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;How is&amp;nbsp;Instructed-Retriever different from the latest&amp;nbsp;Instructed-Retriever-1 a.k.a Knowledge Assistant?&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;Instructed-Retriever is more linear in nature - as in, it does not have any parallel searches, query comparisons etc., However the new Knowledge Assistant can run parallel queries and intelligently combined the evidence output to rank the results to give the best evidence.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;So in terms of quality and model accuracy I would say the new Knowledge Assistant is far better than the older one. However it will be proved eventually on several groundtruth facts validated earlier.&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Sat, 08 Aug 2026 08:01:37 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/databricks-knowledge-assistant/m-p/165136#M1997</guid>
      <dc:creator>suryaprayaga</dc:creator>
      <dc:date>2026-08-08T08:01:37Z</dc:date>
    </item>
    <item>
      <title>Databricks Knowledge Assistant</title>
      <link>https://community.databricks.com/t5/generative-ai/databricks-knowledge-assistant/m-p/165132#M1996</link>
      <description>&lt;P class=""&gt;Hi All,&lt;/P&gt;&lt;P class=""&gt;I've been exploring Knowledge Assistant and trying to understand how it works under the hood. I noticed it's built on the Instructed Retriever architecture, and I recently came across an update called Instructed-Retriever-1, which is supposed to deliver much faster results. Has this already been rolled out to all Databricks accounts, or is it still limited to certain users? Also, curious to understand - what exactly is different about Instructed-Retriever-1 compared to the original Instructed Retriever?&lt;/P&gt;</description>
      <pubDate>Sat, 08 Aug 2026 06:50:08 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/databricks-knowledge-assistant/m-p/165132#M1996</guid>
      <dc:creator>IM_01</dc:creator>
      <dc:date>2026-08-08T06:50:08Z</dc:date>
    </item>
    <item>
      <title>Re: The Open-Weight Revolution: A Game Changer for Our LLM Cost Optimization Odyssey</title>
      <link>https://community.databricks.com/t5/generative-ai/the-open-weight-revolution-a-game-changer-for-our-llm-cost/m-p/165121#M1995</link>
      <description>&lt;P&gt;Great analysis — this mirrors almost exactly what we've been studying in production at a regulated financial institution in Southeast Asia (banking sector).&lt;/P&gt;&lt;P&gt;The "token price is a poor indicator of actual costs" point is the core insight most teams miss. What you're observing with GLM-5.2 is what happens when a model's reasoning architecture aligns well with the task structure: it hits a sweet spot where fewer tokens are needed precisely because it reasons more directly. The cost-efficiency relationship is nonlinear — a model that is 3x cheaper per token but 2x more verbose per task saves you nothing; GLM-5.2 wins on both dimensions simultaneously, which is rare.&lt;/P&gt;&lt;P&gt;A few angles worth adding to your framework:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;1. Self-hosting changes the break-even math at scale&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;At 10K+ tasks/day, the API pricing story shifts materially. A GLM-5.2 deployment on 8×H100s runs roughly $55–70K/month depending on cloud and reservation type. At your $1,280/day API rate scaled to 10K tasks, that is ~$38,400/month — so self-hosting breaks even around month 6, after which marginal inference cost drops 70–80%. For regulated industries with data residency requirements, the compliance value is additive on top of that.&lt;/P&gt;&lt;P&gt;The practical catch: cold-start latency for a 753B MoE model is non-trivial. Plan for warm routing — keep a persistent serving instance for daytime load, use spot capacity for batch overnight jobs.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;2. Model routing on Databricks — practical implementation&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;The tiered routing you describe (Flash → GLM-5.2 → Claude) maps cleanly to Databricks AI Gateway with custom routing logic:&lt;/P&gt;&lt;UL&gt;&lt;LI&gt;Host GLM-5.2 via Mosaic AI Model Serving (or proxy to external GLM API endpoint)&lt;/LI&gt;&lt;LI&gt;Add a lightweight classifier (a 1B BERT-class model is sufficient) as the routing layer — predicts task complexity from prompt features like length, entity density, and instruction type&lt;/LI&gt;&lt;LI&gt;Route to Claude only when classifier confidence + task-type signal crosses a threshold&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;We are seeing 60–70% of enterprise AI tasks land in the "standard" tier in practice. The 15–20% hardest tasks — multi-repository dependency resolution, architecture decisions, ambiguous multi-step reasoning — are where frontier models still earn their premium.&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;3. The regulated-industry sovereignty argument is underweighted&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;Your sovereignty point deserves more emphasis. For banking and financial services it is not just about data residency — it is about auditability. MIT-licensed open weights let you reproduce any inference exactly, version model checkpoints alongside your model risk frameworks, and demonstrate to regulators precisely what the model saw and how it responded. That is structurally impossible with black-box API calls where model versions drift silently.&lt;/P&gt;&lt;P&gt;One open question for anyone running this in production: how are you handling session-level cost attribution across business units? When multiple teams share a single serving endpoint, allocating inference costs back to cost centers is where we see the most operational friction. Would be interested to hear how others are solving this.&lt;/P&gt;</description>
      <pubDate>Sat, 08 Aug 2026 02:30:48 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/the-open-weight-revolution-a-game-changer-for-our-llm-cost/m-p/165121#M1995</guid>
      <dc:creator>DoTA</dc:creator>
      <dc:date>2026-08-08T02:30:48Z</dc:date>
    </item>
    <item>
      <title>Re: How do I query genie space's actual conversation messages? (Genie space API)</title>
      <link>https://community.databricks.com/t5/generative-ai/how-do-i-query-genie-space-s-actual-conversation-messages-genie/m-p/165115#M1994</link>
      <description>&lt;P&gt;Hi Cobs, I hit the same wall building on the Genie API. The key is that content in the message object is the &lt;STRONG&gt;user's&lt;/STRONG&gt; message, not Genie's answer. The actual response lives in the attachments array, and there are several attachment types:&lt;/P&gt;&lt;BLOCKQUOTE&gt;- attachments[].query - the SQL query Genie generated (what you're seeing now)&lt;BR /&gt;- attachments[].text -&amp;nbsp;&lt;STRONG&gt;Genie's natural-language response, including the final summary when available,&lt;/STRONG&gt;&amp;nbsp;this is what the UI renders as the answer&lt;BR /&gt;- attachments[].suggested_questions&amp;nbsp; follow-up suggestions&lt;BR /&gt;- attachments[].viz - visualization (beta, when enabled)&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;So for a message like your example, you'd pull attachments[].text (and fall back to attachments[].query when text is absent, e.g. SQL-only responses). For messages where Genie didn't run a query at all, the text attachment is the only place the answer exists.&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;Two practical tips:&lt;BR /&gt;1. &lt;STRONG&gt;Poll &lt;/STRONG&gt;status until it's COMPLETED before reading attachments, the text attachment isn't guaranteed to be present while Genie is still working (and watch for error.type on failure).&lt;BR /&gt;2. To render a full conversation like the UI, use the &lt;STRONG&gt;List messages&lt;/STRONG&gt; endpoint (GET /api/2.0/genie/spaces/{space_id}/conversations/{conversation_id}/messages) and iterate: user question from content, assistant answer from attachments[].text / attachments[].query.&lt;/BLOCKQUOTE&gt;&lt;BLOCKQUOTE&gt;Dump the full JSON of a known-good message to see all attachment types in your workspace - the docs list them but seeing the real shape helps.&lt;BR /&gt;Hope that unblocks you.&lt;/BLOCKQUOTE&gt;</description>
      <pubDate>Fri, 07 Aug 2026 23:51:02 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/how-do-i-query-genie-space-s-actual-conversation-messages-genie/m-p/165115#M1994</guid>
      <dc:creator>empire_labs</dc:creator>
      <dc:date>2026-08-07T23:51:02Z</dc:date>
    </item>
    <item>
      <title>Re: Building an Agentic HR Front Door on Databricks</title>
      <link>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165107#M1993</link>
      <description>&lt;P&gt;Thanks for your feedback!&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/247096"&gt;@wiliam65jenny&lt;/a&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Fri, 07 Aug 2026 16:17:31 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165107#M1993</guid>
      <dc:creator>Brahmareddy</dc:creator>
      <dc:date>2026-08-07T16:17:31Z</dc:date>
    </item>
    <item>
      <title>Re: Building an Agentic HR Front Door on Databricks</title>
      <link>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165068#M1992</link>
      <description>&lt;P&gt;Hello,&lt;BR /&gt;&lt;BR /&gt;This is a great exploration of where enterprise AI is heading. The focus on moving beyond chat-based assistance toward governed action-taking agents is especially important. I like the emphasis on ontology, memory, tools, and auditability because trust will be the biggest factor in enterprise adoption.&lt;BR /&gt;&lt;BR /&gt;&lt;BR /&gt;Best Regards&lt;/P&gt;</description>
      <pubDate>Fri, 07 Aug 2026 05:25:50 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165068#M1992</guid>
      <dc:creator>wiliam65jenny</dc:creator>
      <dc:date>2026-08-07T05:25:50Z</dc:date>
    </item>
    <item>
      <title>Building an Agentic HR Front Door on Databricks</title>
      <link>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165060#M1991</link>
      <description>&lt;P&gt;I have been exploring one question for the last few months:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;What happens when enterprise AI moves from answering questions to actually completing tasks?&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;That led me to build &lt;STRONG&gt;DwaraOne&lt;/STRONG&gt;, an agentic HR front door using Databricks free edition.&lt;/P&gt;&lt;P&gt;Most HR assistants today can explain a policy or answer a question.&lt;/P&gt;&lt;P&gt;But an employee usually wants more than an answer.&lt;/P&gt;&lt;P&gt;For example, when someone asks for leave, the agent should be able to check the balance, understand team availability, identify possible conflicts, reference the right policy, submit the request, route it to the correct approver, and track what happens next.&lt;/P&gt;&lt;P&gt;That is the shift I wanted to explore.&lt;/P&gt;&lt;P&gt;I built the current preview using &lt;STRONG&gt;Databricks Free Edition&lt;/STRONG&gt;, which made the exercise even more interesting.&lt;/P&gt;&lt;P&gt;The architecture includes:&lt;/P&gt;&lt;P&gt;Unity Catalog for governed data access&lt;/P&gt;&lt;P&gt;Silver and Gold layers with Employee 360 views&lt;/P&gt;&lt;P&gt;SQL Warehouse and Delta for analytics and application data&lt;/P&gt;&lt;P&gt;Genie Spaces for employee, manager, and leader experiences&lt;/P&gt;&lt;P&gt;Lakebase for conversation and workflow state&lt;/P&gt;&lt;P&gt;Governed agent tools for actions&lt;/P&gt;&lt;P&gt;Memory across conversations&lt;/P&gt;&lt;P&gt;Audit logging for agent decisions and actions&lt;/P&gt;&lt;P&gt;One of my biggest learnings was that agent architecture should not start with the chatbot.&lt;/P&gt;&lt;P&gt;It should start with the data model, permissions, context, tools, and auditability.&lt;/P&gt;&lt;P&gt;For me, the important loop became:&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;Ontology → Memory → Tools → Audit&lt;/STRONG&gt;&lt;/P&gt;&lt;P&gt;An agent that can act should also be explainable and accountable.&lt;/P&gt;&lt;P&gt;Another interesting lesson was that Free Edition did not feel like a limitation.&lt;/P&gt;&lt;P&gt;It actually pushed me to think more carefully about governance, semantic design, lineage, and architecture before adding more complexity.&lt;/P&gt;&lt;P&gt;The current preview uses synthetic data and supports scenarios like leave requests, policy guidance, onboarding, offboarding, workforce analytics, manager insights, consent tracking, and agent health monitoring.&lt;/P&gt;&lt;P&gt;I am still improving it and would really value feedback from the Databricks community.&lt;/P&gt;&lt;P&gt;What would you add to this architecture?&lt;/P&gt;&lt;P&gt;And what is the first enterprise task you would trust an agent to complete?&lt;/P&gt;&lt;P&gt;Preview: &lt;A title="DwaraOne - AI Enabled HR Front Door" href="http://dwaraone.online" target="_self"&gt;dwaraone.online&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;div class="video-embed-center video-embed"&gt;&lt;iframe class="embedly-embed" src="https://cdn.embedly.com/widgets/media.html?src=https%3A%2F%2Fwww.youtube.com%2Fembed%2FgqxEJZDykDw%3Ffeature%3Doembed&amp;amp;display_name=YouTube&amp;amp;url=https%3A%2F%2Fwww.youtube.com%2Fwatch%3Fv%3DgqxEJZDykDw&amp;amp;image=https%3A%2F%2Fi.ytimg.com%2Fvi%2FgqxEJZDykDw%2Fhqdefault.jpg&amp;amp;type=text%2Fhtml&amp;amp;schema=youtube" width="600" height="338" scrolling="no" title="Building an Agentic HR Front Door on Databricks" frameborder="0" allow="autoplay; fullscreen; encrypted-media; picture-in-picture" allowfullscreen="true"&gt;&lt;/iframe&gt;&lt;/div&gt;&lt;/P&gt;</description>
      <pubDate>Fri, 07 Aug 2026 03:04:57 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/building-an-agentic-hr-front-door-on-databricks/m-p/165060#M1991</guid>
      <dc:creator>Brahmareddy</dc:creator>
      <dc:date>2026-08-07T03:04:57Z</dc:date>
    </item>
    <item>
      <title>Re: Supervisor Agent or Knowledge assistant not available</title>
      <link>https://community.databricks.com/t5/generative-ai/supervisor-agent-or-knowledge-assistant-not-available/m-p/164878#M1988</link>
      <description>&lt;P&gt;Hi,&lt;/P&gt;
&lt;P&gt;Sorry you're still having issues but glad it seems to be working on the AWS one. It should be working in that region in Azure as far as I can tell. It could be an issue with model availability on Azure, at the point you tried it. If you have an Azure agreement with support you can actually file a support ticket through that and they may be able to advise.&lt;/P&gt;
&lt;P&gt;&lt;BR /&gt;Thanks,&lt;BR /&gt;&lt;BR /&gt;Emma&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Tue, 04 Aug 2026 15:46:04 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/supervisor-agent-or-knowledge-assistant-not-available/m-p/164878#M1988</guid>
      <dc:creator>emma_s</dc:creator>
      <dc:date>2026-08-04T15:46:04Z</dc:date>
    </item>
    <item>
      <title>Re: Supervisor Agent or Knowledge assistant not available</title>
      <link>https://community.databricks.com/t5/generative-ai/supervisor-agent-or-knowledge-assistant-not-available/m-p/164744#M1987</link>
      <description>&lt;UL&gt;&lt;LI&gt;&lt;SPAN&gt;Azure :&amp;nbsp;&amp;nbsp;&lt;/SPAN&gt;&lt;SPAN&gt;swedencentral&amp;nbsp; - everything should be supported&lt;/SPAN&gt;&lt;/LI&gt;&lt;LI&gt;&lt;SPAN&gt;&lt;SPAN&gt;Databricks (native ? - the one you convert to premium after a trial, not triggered from AWS or Azure don't know how you call that deploy option) seems to be running on AWS us-west-2. Also a region where everything is supported.&lt;BR /&gt;&lt;BR /&gt;| just notied that the supervisor agent started working on us-west-2. I did contact sales (As indicated in the error msg) but never got a reply back. perhaps they did something ?&lt;BR /&gt;&lt;/SPAN&gt;&lt;/SPAN&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="tensorwrangler_1-1785764674862.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/29697iE809EE37A1E78A0F/image-size/medium?v=v2&amp;amp;px=400" role="button" title="tensorwrangler_1-1785764674862.png" alt="tensorwrangler_1-1785764674862.png" /&gt;&lt;/span&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;BR /&gt;the one on Azure stil shows me the same error&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;&lt;/UL&gt;&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="tensorwrangler_0-1785764629401.png" style="width: 400px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/29696iE2CC4BA3F683FB5E/image-size/medium?v=v2&amp;amp;px=400" role="button" title="tensorwrangler_0-1785764629401.png" alt="tensorwrangler_0-1785764629401.png" /&gt;&lt;/span&gt;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;SPAN&gt;&lt;BR /&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;</description>
      <pubDate>Mon, 03 Aug 2026 13:44:47 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/supervisor-agent-or-knowledge-assistant-not-available/m-p/164744#M1987</guid>
      <dc:creator>tensorwrangler</dc:creator>
      <dc:date>2026-08-03T13:44:47Z</dc:date>
    </item>
    <item>
      <title>Re: Supervisor Agent or Knowledge assistant not available</title>
      <link>https://community.databricks.com/t5/generative-ai/supervisor-agent-or-knowledge-assistant-not-available/m-p/164710#M1986</link>
      <description>&lt;P&gt;Hi, if you're seeing "Supervisor Agent or Knowledge Assistant not available" across multiple Databricks workspaces (including Databricks-native and Azure), even with preview features enabled and appropriate billing, it's worth checking whether these features are supported in your workspace region and enabled at both the account and workspace levels. In some cases, availability depends on the cloud provider, workspace region, account entitlements, or a phased rollout rather than workspace configuration. You can review the prerequisites here:&lt;/P&gt;&lt;P&gt;&lt;A href="https://docs.databricks.com/aws/en/agents/" target="_blank"&gt;https://docs.databricks.com/aws/en/agents/&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;A href="https://docs.databricks.com/gcp/en/agents/agent-bricks/multi-agent-supervisor" target="_blank"&gt;https://docs.databricks.com/gcp/en/agents/agent-bricks/multi-agent-supervisor&lt;/A&gt;&lt;/P&gt;&lt;P&gt;&lt;A href="https://docs.databricks.com/aws/en/generative-ai/ai-assistant/" target="_blank"&gt;https://docs.databricks.com/aws/en/generative-ai/ai-assistant/&lt;/A&gt;&lt;/P&gt;&lt;P&gt;Could you also share which cloud provider (AWS, Azure, or GCP) and region your workspaces are running in? If everything appears to be configured correctly, this may require Databricks Support to verify feature availability for your account.&lt;/P&gt;</description>
      <pubDate>Mon, 03 Aug 2026 08:42:05 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/supervisor-agent-or-knowledge-assistant-not-available/m-p/164710#M1986</guid>
      <dc:creator>honey_sharma</dc:creator>
      <dc:date>2026-08-03T08:42:05Z</dc:date>
    </item>
    <item>
      <title>Re: Getting error in databricks agent response</title>
      <link>https://community.databricks.com/t5/generative-ai/getting-error-in-databricks-agent-response/m-p/164590#M1984</link>
      <description>&lt;P&gt;Hello! So I am using the built-in code when we initiate the databricks custom app using agent. After tweaking and using AI by Gemini and Claude, I was finally able to start the conversation and receive the response! Sharing the code below for anyone to play with. Of course, this code only returns the first round of conversation and subsequently fails in second turn. I may look into this further, but it seems I may have to flag this to Databricks support as the built-in template should work out of the box regardless it is a reasoning or an instruction model.&amp;nbsp;&lt;/P&gt;&lt;LI-CODE lang="markup"&gt;import json
import logging
from contextlib import AsyncExitStack
from datetime import datetime
from typing import AsyncGenerator

# --- PYDANTIC MONKEY-PATCH FOR REASONING MODELS ---
def _flatten_text_payload(value):
    if value is None:
        return ""
    if isinstance(value, str):
        return value
    if isinstance(value, list):
        parts = []
        for chunk in value:
            if isinstance(chunk, dict):
                if chunk.get("type") in ("text", "output_text"):
                    parts.append(str(chunk.get("text", "")))
                elif "text" in chunk:
                    parts.append(str(chunk.get("text", "")))
            elif isinstance(chunk, str):
                parts.append(chunk)
            else:
                parts.append(str(chunk))
        return "".join(parts)
    if isinstance(value, dict):
        if value.get("type") in ("text", "output_text"):
            return str(value.get("text", ""))
        return str(value.get("text", str(value)))
    return str(value)


def _patch_pydantic_class(cls):
    if cls is None:
        return
    orig_init = cls.__init__
    def patched_init(self, *args, **kwargs):
        if "text" in kwargs:
            kwargs["text"] = _flatten_text_payload(kwargs["text"])
        orig_init(self, *args, **kwargs)
    cls.__init__ = patched_init

    orig_validate = getattr(cls, "model_validate", None)
    if orig_validate:
        @classmethod
        def patched_validate(target_cls, obj, *args, **kwargs):
            if isinstance(obj, dict) and "text" in obj:
                obj["text"] = _flatten_text_payload(obj["text"])
            return orig_validate(obj, *args, **kwargs)
        cls.model_validate = patched_validate


try:
    from openai.types.responses import ResponseOutputText as OpenAIResponseOutputText
    _patch_pydantic_class(OpenAIResponseOutputText)
except ImportError:
    pass

try:
    from mlflow.types.responses import ResponseOutputText as MLflowResponseOutputText
    _patch_pydantic_class(MLflowResponseOutputText)
except ImportError:
    pass
# ---------------------------------------------------

import mlflow
from agents import Agent, Runner, function_tool, set_default_openai_api, set_default_openai_client
from agents.tracing import set_trace_processors
from databricks.sdk import WorkspaceClient
from databricks_openai import AsyncDatabricksOpenAI
from databricks_openai.agents import McpServer
from mlflow.genai.agent_server import invoke, stream
from mlflow.types.responses import (
    ResponsesAgentRequest,
    ResponsesAgentResponse,
    ResponsesAgentStreamEvent,
)

from agent_server.utils import build_mcp_url, get_session_id

logger = logging.getLogger(__name__)

set_default_openai_client(AsyncDatabricksOpenAI())
set_default_openai_api("chat_completions")
set_trace_processors([])
mlflow.openai.autolog()
logging.getLogger("mlflow.utils.autologging_utils").setLevel(logging.ERROR)


def _sanitize_item_dict(dump: dict) -&amp;gt; dict:
    """
    Repair malformed 'text' leaves (some reasoning models return text as a
    nested list/dict instead of a plain string) WITHOUT collapsing the
    container shapes (content / summary) that the Agents SDK's
    chat_completions converter expects to remain lists of parts.

    IMPORTANT: do not flatten `content` or `summary` themselves into plain
    strings -- Converter.items_to_messages() in the openai-agents SDK
    indexes into these as list-of-parts (e.g. content[0]["text"]) when
    reconstructing multi-turn history. Flattening them causes:
    TypeError: string indices must be integers, not 'str'
    on the *second* turn, once history round-trips back through the API.
    """
    # Fix a malformed top-level "text" leaf without touching containers
    if isinstance(dump.get("text"), (list, dict)):
        dump["text"] = _flatten_text_payload(dump["text"])

    # "content" must stay a list-of-parts for the chat_completions converter --
    # only repair malformed "text" leaves inside each part
    content = dump.get("content")
    if isinstance(content, list):
        for part in content:
            if isinstance(part, dict) and isinstance(part.get("text"), (list, dict)):
                part["text"] = _flatten_text_payload(part["text"])
    elif content is not None and not isinstance(content, (str, list)):
        dump["content"] = _flatten_text_payload(content)

    # Reasoning items use "summary": [{"type": "summary_text", "text": ...}]
    summary = dump.get("summary")
    if isinstance(summary, list):
        for part in summary:
            if isinstance(part, dict) and isinstance(part.get("text"), (list, dict)):
                part["text"] = _flatten_text_payload(part["text"])

    return dump


def sanitize_input_messages(input_items):
    clean = []
    for item in input_items:
        dump = item.model_dump() if hasattr(item, "model_dump") else item
        if isinstance(dump, dict):
            clean.append(_sanitize_item_dict(dump))
        else:
            clean.append(dump)
    return clean


@function_tool
def get_current_time() -&amp;gt; str:
    """Get the current date and time."""
    return datetime.now().isoformat()


def create_agent(mcp_servers: list[McpServer] | None = None) -&amp;gt; Agent:
    return Agent(
        name="Agent",
        instructions="You are a helpful assistant.",
        model="databricks-qwen35-122b-a10b",
        tools=[get_current_time],
        mcp_servers=mcp_servers or [],
    )


@invoke()
async def invoke_handler(request: ResponsesAgentRequest) -&amp;gt; ResponsesAgentResponse:
    if session_id := get_session_id(request):
        mlflow.update_current_trace(metadata={"mlflow.trace.session": session_id})

    try:
        async with AsyncExitStack() as stack:
            agent = create_agent()
            messages = sanitize_input_messages(request.input)
            result = await Runner.run(agent, messages)

            cleaned_outputs = []
            for item in result.new_items:
                input_item = item.to_input_item()
                if isinstance(input_item, dict):
                    input_item = _sanitize_item_dict(input_item)
                cleaned_outputs.append(input_item)

            return ResponsesAgentResponse(output=cleaned_outputs)
    except Exception:
        logger.exception("invoke_handler failed")
        raise


@stream()
async def stream_handler(
    request: ResponsesAgentRequest,
) -&amp;gt; AsyncGenerator[ResponsesAgentStreamEvent, None]:
    if session_id := get_session_id(request):
        mlflow.update_current_trace(metadata={"mlflow.trace.session": session_id})

    try:
        async with AsyncExitStack() as stack:
            agent = create_agent()
            messages = sanitize_input_messages(request.input)
            result = Runner.run_streamed(agent, input=messages)

            # Directly stream events safely without process_agent_stream_events
            async for event in result.stream_events():
                delta_val = None
                if hasattr(event, "delta"):
                    delta_val = _flatten_text_payload(event.delta)
                elif hasattr(event, "data") and hasattr(event.data, "delta"):
                    delta_val = _flatten_text_payload(event.data.delta)
                elif hasattr(event, "item") and hasattr(event.item, "text"):
                    delta_val = _flatten_text_payload(event.item.text)

                if delta_val:
                    yield ResponsesAgentStreamEvent(delta=delta_val)
    except Exception:
        logger.exception("stream_handler failed")
        raise&lt;/LI-CODE&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Fri, 31 Jul 2026 08:39:22 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/getting-error-in-databricks-agent-response/m-p/164590#M1984</guid>
      <dc:creator>ajaygshah</dc:creator>
      <dc:date>2026-07-31T08:39:22Z</dc:date>
    </item>
    <item>
      <title>Re: Billing structure and LLM usage costs for no-code agents</title>
      <link>https://community.databricks.com/t5/generative-ai/billing-structure-and-llm-usage-costs-for-no-code-agents/m-p/164496#M1983</link>
      <description>&lt;P&gt;Hi&amp;nbsp;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/166789"&gt;@r_w_&lt;/a&gt;&amp;nbsp;!&lt;/P&gt;&lt;P&gt;Estimating cost and understanding the billing model for No-Code Agents and Genie in Databricks comes down to how DBUs (Databricks Compute Units) and Token consumption are metered across three core layers:&lt;/P&gt;&lt;P&gt;1. Overview of Where Billing Occurs&lt;BR /&gt;There is no standalone "base fee" just for maintaining an Agent definition. You are billed purely based on consumption across these items:&lt;/P&gt;&lt;P&gt;Orchestration &amp;amp; Agent Execution: Charged under Serverless Real-Time Inference DBUs (Model Serving / Agent Hosting) for the compute hosting the agent logic and Supervisor routing.&lt;/P&gt;&lt;P&gt;SQL Warehouse Usage (for Genie Agent): When Genie generates and runs SQL queries against your Lakehouse, you are billed for the underlying Databricks SQL Warehouse DBUs (Serverless or Pro).&lt;/P&gt;&lt;P&gt;LLM Inference Costs: Billed via Foundation Model API (FMAPI) on a Pay-per-token basis (input/output tokens) or Provisioned Throughput if configured.&lt;/P&gt;&lt;P&gt;2. LLM Usage Costs (Included or Billed Separately?)&lt;BR /&gt;LLM costs are billed separately based on token usage, tracked via Foundation Model API / External Models Gateway.&lt;/P&gt;&lt;P&gt;When a Supervisor Agent or Genie routes a prompt to a managed LLM, the tokens consumed (prompt + completion) are metered through FMAPI.&lt;/P&gt;&lt;P&gt;The execution runtime (Serverless Inference DBUs) and the LLM token usage (FMAPI) will appear as distinct line items in your Azure/AWS/GCP Databricks bill.&lt;/P&gt;&lt;P&gt;3. Which LLM is Used &amp;amp; Can It Be Changed?&lt;BR /&gt;Default LLMs: By default, Genie and Agent Bricks leverage Databricks' optimized foundation models (such as DBRX, Llama 3, or partner models integrated natively).&lt;/P&gt;&lt;P&gt;Can you change it? Yes. You can route and customize the underlying models through the Databricks AI Gateway / Model Serving endpoints. This allows you to point your agents to different foundation models (e.g., Anthropic Claude, OpenAI GPT-4o, Llama 3, or custom fine-tuned models) depending on your accuracy and cost requirements.&lt;/P&gt;&lt;P&gt;Official Documentation for Cost Estimation:&lt;BR /&gt;Databricks Pricing Page (Model Serving &amp;amp; FMAPI): &lt;A href="https://www.databricks.com/product/pricing" target="_blank"&gt;https://www.databricks.com/product/pricing&lt;/A&gt;&lt;/P&gt;&lt;P&gt;Databricks Genie Billing &amp;amp; Architecture:&amp;nbsp;&lt;A href="https://docs.databricks.com/en/genie/index.html" target="_self"&gt;https://docs.databricks.com/en/genie/index.html&lt;/A&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;&lt;STRONG&gt;If my answer was helpful, please consider marking it as accepted solution!&lt;/STRONG&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 30 Jul 2026 13:04:21 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/billing-structure-and-llm-usage-costs-for-no-code-agents/m-p/164496#M1983</guid>
      <dc:creator>GabFernandes</dc:creator>
      <dc:date>2026-07-30T13:04:21Z</dc:date>
    </item>
    <item>
      <title>Re: Genie task running for over 20 hours.</title>
      <link>https://community.databricks.com/t5/generative-ai/genie-task-running-for-over-20-hours/m-p/164477#M1990</link>
      <description>&lt;P&gt;Hi egs,&lt;/P&gt;&lt;DIV&gt;&lt;P&gt;You can delete the notebook or chart to get rid of the hung Genie task.&amp;nbsp;Since Genie tasks are generating and executing read only exploratory queries against the SQL Warehouse, they generally don't modify the underlying Delta tables. Deleting the UI artifact must simply tear down the client session and abandon the execution state.&lt;/P&gt;&lt;P&gt;Before you delete it, you can first confirm whether this is just a UI desync. You can do a refresh of the workspace page or try closing and reopening the tab as it often clears a stuck visual state if the query has actually already finished or timed out on the backend.&lt;/P&gt;&lt;P&gt;If the task is genuinely still hanging in the UI after 20 hours, go ahead and delete the notebook or chart.&amp;nbsp;To prevent it from hanging again, you can instruct Genie use a much smaller date range or add&amp;nbsp;LIMIT clause in its queries.&lt;/P&gt;&lt;/DIV&gt;</description>
      <pubDate>Thu, 30 Jul 2026 11:08:04 GMT</pubDate>
      <guid>https://community.databricks.com/t5/generative-ai/genie-task-running-for-over-20-hours/m-p/164477#M1990</guid>
      <dc:creator>balajij8</dc:creator>
      <dc:date>2026-07-30T11:08:04Z</dc:date>
    </item>
  </channel>
</rss>

