<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>article Building a Production LangGraph Agent on Databricks - NorthStar Brand Copilot in Technical Blog</title>
    <link>https://community.databricks.com/t5/technical-blog/building-a-production-langgraph-agent-on-databricks-northstar/ba-p/162634</link>
    <description>&lt;P&gt;&lt;SPAN&gt;Building a Production LangGraph Agent on Databricks: The NorthStar Brand Copilot&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;How we built and deployed an end-to-end CPG AI agent on Databricks using LangGraph, MCP servers, Lakebase memory, Genie, Vector Search, and MLflow and the best practices we learned along the way.&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;TL;DR&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;We built &lt;/SPAN&gt;&lt;STRONG&gt;NorthStar Brand Copilot&lt;/STRONG&gt;&lt;SPAN&gt;, an AI assistant for brand managers and field-sales reps at a fictional multi-category CPG company (Snacks, Beverages, Personal Care). It's a &lt;/SPAN&gt;&lt;STRONG&gt;LangGraph custom agent&lt;/STRONG&gt;&lt;SPAN&gt; that routes each question to the right &lt;/SPAN&gt;&lt;STRONG&gt;Databricks-native&lt;/STRONG&gt;&lt;SPAN&gt; capability:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Genie&lt;/STRONG&gt;&lt;SPAN&gt; (NatualLanguage→SQL) for the numbers — sell-in/sell-out, trade-promotion ROI, inventory, market share&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;AI Search&lt;/STRONG&gt;&lt;SPAN&gt;(RAG) for the documents — product specs, allergens, consumer reviews, brand guidelines, the promo playbook&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Lakebase&lt;/STRONG&gt;&lt;SPAN&gt; (long-term memory) for the decisions, For ex: "remember we cut BOGO (Buy One Get One) at Walgreens," recalled across sessions. This provides a persistent, semantic context layer, ensuring the agent doesn't process queries in isolation but evolves its understanding based on past decisions and specific user history.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;LangGraph&lt;/STRONG&gt;&lt;SPAN&gt; for agent orchestration. While simpler frameworks (like standard LangChain chains or monolithic agent patterns) are easier to start with, LangGraph was chosen for its superior ability to handle cyclical, complex workflows and granular state management. It allows the agent to reason, refine its tool-calling strategy iteratively, and maintain persistent state, which is essential for enterprise-grade reliability compared to more rigid, linear alternatives.&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;While LangGraph is ideal for cyclical and stateful workflows, other orchestration frameworks offer different strengths, Below are some other alternatives:&lt;/SPAN&gt;&lt;/LI&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="2"&gt;&lt;STRONG&gt;LangChain (Core/Chains):&lt;/STRONG&gt;&lt;SPAN&gt; Best for linear, straightforward pipelines where complex looping or persistent state management is not the primary requirement. It offers a vast ecosystem of integrations.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="2"&gt;&lt;STRONG&gt;LlamaIndex:&lt;/STRONG&gt;&lt;SPAN&gt; Excels in data-heavy applications, providing highly optimized query engines and ingestion pipelines for RAG-centric workflows.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="2"&gt;&lt;STRONG&gt;DSPy:&lt;/STRONG&gt;&lt;SPAN&gt; Focuses on production reliability through a declarative approach, programmatically optimizing prompts and module assembly to improve performance systematically rather than through manual tuning.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;&lt;STRONG&gt;Note&lt;/STRONG&gt;: The CPG organization and data used in this blog post are fictional and intended for&amp;nbsp;&lt;/SPAN&gt;&lt;/I&gt;&lt;I&gt;&lt;SPAN&gt;demonstration purposes only.&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Full Source Code Repository&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;The full source code, including synthetic data generation, setup scripts and deployment automation, is available in the &lt;/SPAN&gt;&lt;A title="code_repo" href="https://github.com/databricks-solutions/databricks-blogposts/tree/main/2026-07-cpg-brand-copilot-agent" target="_self"&gt;source_code_repo&lt;/A&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;The whole thing runs &lt;/SPAN&gt;&lt;STRONG&gt;inside a single Databricks App&lt;/STRONG&gt;&lt;SPAN&gt;, deployed with &lt;/SPAN&gt;&lt;STRONG&gt;Declarative Automation Bundles (DABs)&lt;/STRONG&gt;&lt;SPAN&gt;, Agent powered by &lt;/SPAN&gt;&lt;STRONG&gt;Claude Sonnet 4.5&lt;/STRONG&gt;&lt;SPAN&gt;, and observed/evaluated with &lt;/SPAN&gt;&lt;STRONG&gt;MLflow&lt;/STRONG&gt;&lt;SPAN&gt;. This post walks through the architecture and the practices that make it production-grade.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Architecture Flow:&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_0-1783743763549.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/28858iF39347B834FDBFA5/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_0-1783743763549.png" alt="SashankKotta_0-1783743763549.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Why this architecture:&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;A few years ago, deploying an agent on Databricks meant logging a model to Unity Catalog and starting up a &lt;/SPAN&gt;&lt;STRONG&gt;Model Serving&lt;/STRONG&gt;&lt;SPAN&gt; endpoint. That still works, but the &lt;/SPAN&gt;&lt;STRONG&gt;current recommended pattern&lt;/STRONG&gt;&lt;SPAN&gt; is to run the agent &lt;/SPAN&gt;&lt;STRONG&gt;inside a Databricks App&lt;/STRONG&gt;&lt;SPAN&gt; as an MLflow &lt;/SPAN&gt;&lt;SPAN&gt;ResponsesAgent&lt;/SPAN&gt;&lt;SPAN&gt; served by &lt;/SPAN&gt;&lt;SPAN&gt;AgentServer&lt;/SPAN&gt;&lt;SPAN&gt;. We deliberately chose this newer path because it gives you:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;One artifact, one URL, one deploy:&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;The UI, the API (&lt;/SPAN&gt;&lt;SPAN&gt;/invocations&lt;/SPAN&gt;&lt;SPAN&gt;), and the agent loop all live in the same app.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Native identity &amp;amp; governance:&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;The app runs as a &lt;/SPAN&gt;&lt;STRONG&gt;service principal&lt;/STRONG&gt;&lt;SPAN&gt;; every tool call is authorized through Unity Catalog and resource grants, no long-lived tokens baked into code.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;First-class observability:&lt;/STRONG&gt;&lt;SPAN&gt; MLflow tracing is wired in with a single autolog line. Agent Evaluation is available too, but it is not automatic, it requires you to define an eval dataset, choose scorers, add the databricks-agents dependency, and set quality thresholds.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;Everything governed by Unity Catalog · traces + eval in MLflow&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;The LangGraph custom agent&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;The heart of this project is &lt;/SPAN&gt;&lt;STRONG&gt;agent_server/agent.py&lt;/STRONG&gt;&lt;SPAN&gt;. It's intentionally small, the framework does the heavy lifting. Here's the shape of it:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;from databricks_langchain import ChatDatabricks
from langchain.agents import create_agent
from mlflow.genai.agent_server import invoke, stream
mlflow.langchain.autolog()                       # full tracing, one line
@stream()
async def stream_handler(request: ResponsesAgentRequest):
    user_messages = to_chat_completions_input([i.model_dump() for i in request.input])
    messages = {"messages": [{"role": "system", "content": AGENT_INSTRUCTIONS}] + user_messages}
    tools = await _build_tools(sp_workspace_client)        # Genie + Vector Search (MCP) + time
    # ... open Lakebase memory store, add memory tools ...
    agent = create_agent(tools=tools, model=ChatDatabricks(endpoint=MODEL_ENDPOINT))
    async for event in process_agent_astream_events(agent.astream(...)):
        yield event&lt;/LI-CODE&gt;
&lt;H3&gt;&lt;STRONG&gt;Best practice 1: Let the model route; describe tools well&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;We don't hand-code a supervisor graph with hard-wired branches. Instead we give a single &lt;/SPAN&gt;&lt;STRONG&gt;create_agent&lt;/STRONG&gt;&lt;SPAN&gt; a well-described tool set and a &lt;/SPAN&gt;&lt;STRONG&gt;routing-oriented system prompt&lt;/STRONG&gt;&lt;SPAN&gt;, and let &lt;/SPAN&gt;&lt;STRONG&gt;Claude&lt;/STRONG&gt;&lt;SPAN&gt; decide which tool(s) to call. The prompt is explicit about &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;when&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; to use each capability:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;"Route quantitative questions to Genie. Route document-based or qualitative questions to Vector Search. For queries requiring both, call Genie first to retrieve numbers, then query AI Search(Formerly Vector Search) for guidance. Always cite sources. Never invent numbers."&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;This is the single highest-leverage piece of the agent. Tool &lt;/SPAN&gt;&lt;STRONG&gt;descriptions&lt;/STRONG&gt;&lt;SPAN&gt; and the system prompt are your routing logic. Invest there before you reach for custom graph code.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Gotcha:&lt;/STRONG&gt; &lt;SPAN&gt;create_agent&lt;/SPAN&gt;&lt;SPAN&gt; does &lt;/SPAN&gt;&lt;STRONG&gt;not&lt;/STRONG&gt;&lt;SPAN&gt; accept a &lt;/SPAN&gt;&lt;SPAN&gt;prompt =&lt;/SPAN&gt;&lt;SPAN&gt; argument the way older &lt;/SPAN&gt;&lt;SPAN&gt;create_react_agent&lt;/SPAN&gt;&lt;SPAN&gt; examples do. Prepend your instructions as a &lt;/SPAN&gt;&lt;STRONG&gt;system message&lt;/STRONG&gt;&lt;SPAN&gt; in the input state instead.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Best practice 2: Build your tools defensively (i.e. graceful degradation)&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;Make sure the _build_tools&lt;/SPAN&gt;&lt;SPAN&gt; always returns &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;something&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; usable, even if a backend is down or a resource was deleted:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;async def _build_tools(workspace_client):
    tools = [get_current_time]
    try:
        mcp_client = init_mcp_client(workspace_client)
        tools.extend(_stringify_tool(t) for t in await mcp_client.get_tools())
    except Exception:
        logger.warning("Failed to fetch MCP tools; continuing.", exc_info=True)
    return tools&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;This mattered in practice: When the Vector Search index was later deleted / Not available in one workspace, the agent kept working, it just lost the RAG path instead of crashing. &lt;/SPAN&gt;&lt;STRONG&gt;An agent that degrades is better than an agent returns a 500 error.&lt;/STRONG&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Best practice 3: Normalize tool output for your LLM provider&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;A subtle but critical fix. Databricks-managed MCP tools (Genie especially) return &lt;/SPAN&gt;&lt;STRONG&gt;structured content blocks&lt;/STRONG&gt;&lt;SPAN&gt; that include an &lt;/SPAN&gt;&lt;SPAN&gt;id&lt;/SPAN&gt;&lt;SPAN&gt; field. The Claude endpoint rejects that field in &lt;/SPAN&gt;&lt;SPAN&gt;tool_result&lt;/SPAN&gt;&lt;SPAN&gt; content with:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;&lt;FONT face="courier new,courier"&gt;tool_result.content.0.text.id:&lt;/FONT&gt; Extra inputs are not permitted&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;The fix is a thin wrapper that coerces every MCP tool's output to a plain string:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;def _stringify_tool(t: StructuredTool) -&amp;gt; StructuredTool:
    async def _wrapped(**kwargs):
        out = await t.ainvoke(kwargs)
        return out if isinstance(out, str) else json.dumps(out, default=str)
    return StructuredTool(
        name=t.name, description=t.description, args_schema=t.args_schema, coroutine=_wrapped
    )&lt;/LI-CODE&gt;
&lt;P&gt;&lt;STRONG&gt;Lesson:&lt;/STRONG&gt;&lt;SPAN&gt; when you bridge MCP tools to a specific model endpoint, validate the tool-result schema. Coercing to plain text is a safe default.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Best practice 4: Pin your model and centralize config&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;The model endpoint, Genie space, AI Search(Formerly Vector Search) catalog/schema, Lakebase instance, and embedding config are all &lt;/SPAN&gt;&lt;STRONG&gt;environment variables with sensible defaults&lt;/STRONG&gt;&lt;SPAN&gt;, set in &lt;/SPAN&gt;&lt;SPAN&gt;databricks.yml&lt;/SPAN&gt;&lt;SPAN&gt; / &lt;/SPAN&gt;&lt;SPAN&gt;app.yaml&lt;/SPAN&gt;&lt;SPAN&gt;. We pin a specific, capable model — &lt;/SPAN&gt;&lt;SPAN&gt;databricks-claude-sonnet-4-5,&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;rather than a floating alias. This keeps behavior reproducible across the multiple workspaces if deployed to.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;MCP servers: the integration layer&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Instead of writing bespoke API clients for Genie and AI Search(Formerly Vectory Search), the agent consumes them as &lt;/SPAN&gt;&lt;STRONG&gt;MCP (Model Context Protocol) servers&lt;/STRONG&gt;&lt;SPAN&gt;. Databricks exposes managed MCP endpoints for its services, and &lt;/SPAN&gt;&lt;SPAN&gt;databricks-langchain&lt;/SPAN&gt;&lt;SPAN&gt; gives you a client that turns them into LangChain tools:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;from databricks_langchain import DatabricksMCPServer, DatabricksMultiServerMCPClient
def init_mcp_client(workspace_client):
    host = get_databricks_host_from_env()
    return DatabricksMultiServerMCPClient([
        DatabricksMCPServer(
            name="genie",
            url=f"{host}/api/2.0/mcp/genie/{GENIE_SPACE_ID}",
            workspace_client=workspace_client,
        ),
        DatabricksMCPServer(
            name="vector-search",
        url=f"{host}/api/2.0/mcp/vector-search/{VS_CATALOG}/{VS_SCHEMA}",
            workspace_client=workspace_client,
        ),
    ])
tools = await mcp_client.get_tools()&lt;/LI-CODE&gt;
&lt;H3&gt;&lt;SPAN&gt;Why MCP is the right call&lt;/SPAN&gt;&lt;/H3&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Managed endpoints, no glue code.&lt;/STRONG&gt;&lt;SPAN&gt;&amp;nbsp;A Genie space becomes a tool at /api/2.0/mcp/genie/&amp;lt;space_id&amp;gt;; an AI Search index is reached at /api/2.0/mcp/ai-search/&amp;lt;catalog&amp;gt;/&amp;lt;schema&amp;gt;/&amp;lt;index_name&amp;gt;, and the schema-level path /api/2.0/mcp/ai-search/&amp;lt;catalog&amp;gt;/&amp;lt;schema&amp;gt; exposes every index in that schema as a tool. The legacy /api/2.0/mcp/vector-search/ prefix still works for backward compatibility. No SDK wiring per service.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Identity flows through.&lt;/STRONG&gt;&lt;SPAN&gt; The MCP client uses the app's &lt;/SPAN&gt;&lt;SPAN&gt;WorkspaceClient&lt;/SPAN&gt;&lt;SPAN&gt;, so calls execute as the app's service principal and are governed by the grants you declared.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Interoperable.&lt;/STRONG&gt;&lt;SPAN&gt; The same MCP pattern extends to &lt;/SPAN&gt;&lt;STRONG&gt;custom MCP servers&lt;/STRONG&gt;&lt;SPAN&gt; (Databricks Apps named &lt;/SPAN&gt;&lt;SPAN&gt;mcp-*&lt;/SPAN&gt;&lt;SPAN&gt;) and external/UC-connection MCP servers, declare the app as an &lt;/SPAN&gt;&lt;SPAN&gt;app&lt;/SPAN&gt;&lt;SPAN&gt; resource in &lt;/SPAN&gt;&lt;SPAN&gt;databricks.yml&lt;/SPAN&gt;&lt;SPAN&gt; and the bundle grants &lt;/SPAN&gt;&lt;SPAN&gt;CAN_USE&lt;/SPAN&gt;&lt;SPAN&gt; on deployment (requires CLI v0.298.0+).&lt;/SPAN&gt;&lt;SPAN&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;I&gt;Note: &lt;/I&gt;&lt;/STRONG&gt;&lt;I&gt;&lt;SPAN&gt;We selected MCP as our preferred standard because it offers a unified, extensible protocol for tool integration. While Genie APIs or standard UC functions could handle specific tasks, MCP provides a future-proof framework. We anticipate expanding the Copilot's capabilities with a broader array of tools and external services, and MCP’s standardized interface significantly simplifies this long-term extensibility.&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt;&lt;BR /&gt;&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;There are three tool types worth knowing on the platform:&amp;nbsp;&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Unity Catalog function tools&lt;/STRONG&gt;&lt;SPAN&gt; (governed SQL UDFs)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;agent code tools&lt;/STRONG&gt;&lt;SPAN&gt; (defined inline, for low-latency REST calls — our &lt;/SPAN&gt;&lt;SPAN&gt;get_current_time&lt;/SPAN&gt;&lt;SPAN&gt; is one)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;MCP tools&lt;/STRONG&gt;&lt;SPAN&gt; (the interoperable path we lean on here).&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;&lt;SPAN&gt;Agent memory: long-term recall with Lakebase&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;This is what turns a chatbot into a &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;copilot&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt;. NorthStar remembers decisions ("we decided to cut BOGO(Buy One Get One) at Walgreens"), flags action items, and recalls them later, even in a brand-new session.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Databricks recognizes two kinds of memory:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Type&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Use case&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Backing&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Key&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Short-term&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;History within one session&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;AsyncCheckpointSaver&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;thread_id&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Long-term&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Facts that persist across sessions&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;AsyncDatabricksStore&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;user_id&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;NorthStar uses &lt;/SPAN&gt;&lt;STRONG&gt;long-term&lt;/STRONG&gt;&lt;SPAN&gt; memory on &lt;/SPAN&gt;&lt;STRONG&gt;Lakebase&lt;/STRONG&gt;&lt;SPAN&gt; (Databricks' managed Postgres) via &lt;/SPAN&gt;&lt;SPAN&gt;AsyncDatabricksStore&lt;/SPAN&gt;&lt;SPAN&gt;, with semantic search powered by the same &lt;/SPAN&gt;&lt;SPAN&gt;databricks-gte-large-en&lt;/SPAN&gt;&lt;SPAN&gt; embedding endpoint used elsewhere.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;SPAN&gt;How it's wired&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;The store is opened as an async context manager and threaded into the agent through &lt;/SPAN&gt;&lt;SPAN&gt;RunnableConfig&lt;/SPAN&gt;&lt;SPAN&gt;:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;async with store_cm as store:
    await store.setup()                  # idempotent; creates store tables on first use
    tools = tools + memory_tools()
    agent = create_agent(tools=tools, model=ChatDatabricks(endpoint=MODEL_ENDPOINT))
    config = {"configurable": {"user_id": user_id, "store": store}}
    async for event in process_agent_astream_events(agent.astream(input=messages, config=config, ...)):
        yield event&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;Memory is exposed to the model as &lt;/SPAN&gt;&lt;STRONG&gt;three tools&lt;/STRONG&gt;,&amp;nbsp;returned from a factory function&lt;STRONG&gt;:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;&lt;SPAN&gt;get_user_memory&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN&gt;save_user_memory&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI&gt;&lt;SPAN&gt;delete_user_memory&lt;/SPAN&gt;&lt;SPAN&gt;&amp;nbsp;&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;The factory pattern lets each tool close over nothing but &lt;/SPAN&gt;&lt;SPAN&gt;RunnableConfig&lt;/SPAN&gt;&lt;SPAN&gt;, from which it pulls the &lt;/SPAN&gt;&lt;SPAN&gt;store&lt;/SPAN&gt;&lt;SPAN&gt; and &lt;/SPAN&gt;&lt;SPAN&gt;user_id&lt;/SPAN&gt;&lt;SPAN&gt;. Memories are namespaced per user: &lt;/SPAN&gt;&lt;SPAN&gt;("user_memories", user_id.replace(".", "-"))&lt;/SPAN&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;SPAN&gt;Memory best practices we followed&lt;/SPAN&gt;&lt;/H3&gt;
&lt;OL&gt;
&lt;LI&gt;&lt;STRONG&gt;Extract &lt;/STRONG&gt;user_id&lt;STRONG&gt; explicitly, return &lt;/STRONG&gt;None&lt;STRONG&gt; on miss.&lt;/STRONG&gt; get_user_id() derives identity from the trusted Databricks Apps OBO context (request.context.user_id) first, and falls back to custom_inputs.user_id only for direct API calls that carry no authenticated user. When both are present they must match — a custom_inputs.user_id that conflicts with the authenticated OBO identity is rejected — so a caller can never read or write another user's memories by supplying someone else's ID. It returns None rather than a default so the &lt;EM&gt;caller&lt;/EM&gt; decides the fallback, never silently write one user's memory under another's namespace.&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Pass the store via config, not as a parameter.&lt;/STRONG&gt;&lt;SPAN&gt; Tools read &lt;/SPAN&gt;&lt;SPAN&gt;config.get("configurable", {}).get("store")&lt;/SPAN&gt;&lt;SPAN&gt;. This keeps tool signatures clean and LLM-friendly.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Validate the LLM's JSON before persisting.&lt;/STRONG&gt; &lt;SPAN&gt;save_user_memory&lt;/SPAN&gt; &lt;SPAN&gt;json.loads&lt;/SPAN&gt;&lt;SPAN&gt; the payload and rejects non-objects, the model &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;will&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; occasionally hand you malformed JSON.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Check &lt;/STRONG&gt;&lt;SPAN&gt;user_id&lt;/SPAN&gt;&lt;STRONG&gt; and &lt;/STRONG&gt;&lt;SPAN&gt;store&lt;/SPAN&gt;&lt;STRONG&gt; separately, with distinct messages.&lt;/STRONG&gt;&lt;SPAN&gt; Easier to debug "no user_id" vs. "store not configured."&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;setup()&lt;/SPAN&gt;&lt;STRONG&gt; is idempotent — call it on startup.&lt;/STRONG&gt;&lt;SPAN&gt; It creates the store tables on first use; safe to run every time.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;STRONG&gt;Two deployment gotchas, both real:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;The &lt;/SPAN&gt;&lt;SPAN&gt;value_from: "database"&lt;/SPAN&gt;&lt;SPAN&gt; binding in &lt;/SPAN&gt;&lt;SPAN&gt;app.yaml&lt;/SPAN&gt;&lt;SPAN&gt; resolves to the Lakebase &lt;/SPAN&gt;&lt;STRONG&gt;hostname&lt;/STRONG&gt;&lt;SPAN&gt;, not the instance name. &lt;/SPAN&gt;&lt;SPAN&gt;AsyncDatabricksStore&lt;/SPAN&gt;&lt;SPAN&gt; wants the instance name, so we call &lt;/SPAN&gt;&lt;SPAN&gt;resolve_lakebase_instance_name()&lt;/SPAN&gt;&lt;SPAN&gt; (it lists DB instances and matches the rw/ro DNS) before constructing the store. Skip this and you get &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;"Unable to resolve Lakebase provisioned instances."&lt;/SPAN&gt;&lt;/I&gt;&lt;/LI&gt;
&lt;LI&gt;Direct /invocations calls &lt;STRONG&gt;must&lt;/STRONG&gt; pass custom_inputs.user_id explicitly — the chat UI / OBO path supplies it automatically, but raw API calls don't. Because a raw caller could supply any value here, restrict direct /invocations access to trusted callers and treat the OBO context as the authoritative identity (any custom_inputs.user_id that conflicts with it is rejected). Without any identity the agent correctly refuses ("requires user identification").&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;Just like tool-building, memory &lt;/SPAN&gt;&lt;STRONG&gt;degrades gracefully&lt;/STRONG&gt;&lt;SPAN&gt; — if &lt;/SPAN&gt;&lt;SPAN&gt;databricks-langchain[memory]&lt;/SPAN&gt;&lt;SPAN&gt; isn't installed or Lakebase is unreachable, the agent drops the memory tools and keeps answering.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Lakebase: Deep Dive and Freshness&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Beyond serving as a persistent key-value store, Lakebase acts as a durable, managed storage layer that allows agents to maintain state beyond short-lived session histories. By leveraging managed Postgres, it provides a consistent, reliable backend that ensures agent memory remains available even when compute clusters spin down or sessions expire, effectively turning stateless LLM calls into a continuous, learning assistant.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Handling Stale Data&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;As with any long-term memory system, stale or outdated information can degrade agent performance. To maintain relevance, implement the following strategies:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="4,0,1,0"&gt;Timestamps &amp;amp; Versioning:&lt;/STRONG&gt; Always include a 'last_updated' timestamp with stored facts. When the agent retrieves memory, the logic should filter for freshness, potentially ignoring facts older than a defined threshold.&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Explicit Update Patterns:&lt;/STRONG&gt;&lt;SPAN&gt; Rather than simply appending new memories, utilize the &lt;/SPAN&gt;&lt;SPAN&gt;save_user_memory&lt;/SPAN&gt;&lt;SPAN&gt; tool to perform 'upserts' (update-or-insert). Overwriting specific keys with the latest information ensures the agent always prioritizes the most current context.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Periodic Pruning:&lt;/STRONG&gt;&lt;SPAN&gt; Implement scheduled cleanup jobs that periodically scan the memory store for expired records or deprecated information, ensuring the Lakebase instance remains performant and focused on relevant data.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Contextual Validation:&lt;/STRONG&gt;&lt;SPAN&gt; When the agent retrieves a memory, prompt it to verify if the information is still applicable based on the current user request, adding an extra layer of human-in-the-loop or logical validation.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;&lt;SPAN&gt;Databricks features, end to end&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;The demo is a tour of the platform's GenAI surface area:&lt;/SPAN&gt;&lt;/P&gt;
&lt;TABLE&gt;
&lt;TBODY&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Capability&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Databricks feature&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Role in the agent&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Analytics (NL→SQL)&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Genie space&lt;/STRONG&gt;&lt;SPAN&gt; over 7 governed Delta tables&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Quantitative questions: ROI, sell-through, inventory&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Document insights (RAG)&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Vector Search&lt;/STRONG&gt;&lt;SPAN&gt; (Delta-sync index, &lt;/SPAN&gt;&lt;SPAN&gt;databricks-gte-large-en&lt;/SPAN&gt;&lt;SPAN&gt;)&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Specs, allergens, reviews, playbook, briefs&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Long-term memory&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Lakebase&lt;/STRONG&gt;&lt;SPAN&gt; (managed Postgres) + &lt;/SPAN&gt;&lt;SPAN&gt;AsyncDatabricksStore&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Persist &amp;amp; recall decisions across sessions&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;LLM&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Foundation Model API&lt;/STRONG&gt;&lt;SPAN&gt; — &lt;/SPAN&gt;&lt;SPAN&gt;databricks-claude-sonnet-4-5&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Reasoning + routing&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Governance&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Unity Catalog&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Every table, index, and function authorized via grants&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Hosting&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Databricks Apps&lt;/STRONG&gt;&lt;SPAN&gt; (FastAPI + custom SPA)&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;One app: Dashboard tab + Assistant tab&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Packaging&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;Automation Bundles (DABs)&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Declarative deploy + resource permissions&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Observability&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;MLflow Tracing&lt;/STRONG&gt;&lt;SPAN&gt; (autolog)&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Per-request span tree: routing → tool → LLM&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;TR&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Quality&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;STRONG&gt;MLflow Agent Evaluation&lt;/STRONG&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;TD&gt;
&lt;P&gt;&lt;SPAN&gt;Scored on Correctness / Relevance / Safety&lt;/SPAN&gt;&lt;/P&gt;
&lt;/TD&gt;
&lt;/TR&gt;
&lt;/TBODY&gt;
&lt;/TABLE&gt;
&lt;P&gt;&lt;SPAN&gt;A nice touch: the app serves a &lt;/SPAN&gt;&lt;STRONG&gt;custom two-tab SPA&lt;/STRONG&gt;&lt;SPAN&gt; (vanilla JS + Chart.js) from the same FastAPI process — a &lt;/SPAN&gt;&lt;STRONG&gt;Dashboard&lt;/STRONG&gt;&lt;SPAN&gt; tab backed by a &lt;/SPAN&gt;&lt;SPAN&gt;/api/analytics&lt;/SPAN&gt;&lt;SPAN&gt; SQL endpoint, and an &lt;/SPAN&gt;&lt;STRONG&gt;Assistant&lt;/STRONG&gt;&lt;SPAN&gt; tab that calls &lt;/SPAN&gt;&lt;SPAN&gt;/invocations&lt;/SPAN&gt;&lt;SPAN&gt;. Running the app &lt;/SPAN&gt;&lt;STRONG&gt;backend-only&lt;/STRONG&gt;&lt;SPAN&gt; (&lt;/SPAN&gt;&lt;SPAN&gt;command: ["uv", "run", "start-server"]&lt;/SPAN&gt;&lt;SPAN&gt;) means the FastAPI app serves both the UI and the agent — no separate Node/Next.js frontend to build.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Best Practice 5 — Production Safeguards: AI Gateway and Guardrails&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;When moving from prototype to production, ensuring stability and safety is non-negotiable.&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;AI Gateway:&lt;/STRONG&gt;&lt;SPAN&gt; We leverage the Model Serving AI Gateway to manage costs and reliability. This provides built-in rate limiting, caching for repeated prompts, and centralized API key management, ensuring our calls to the model are robust and compliant with corporate policies.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Guardrails:&lt;/STRONG&gt;&lt;SPAN&gt; For production deployments, we recommend wrapping your agent in a guardrail layer that inspects inputs and outputs before and after they reach the model. This layer is critical for establishing operational rules, such as screening for unauthorized PII disclosure, enforcing topic alignment, and detecting prompt-injection attempts. These guardrails can be implemented at the application layer or via dedicated safety models and should be treated as a first-class part of your agent's request/response path to ensure consistent reliability and security.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;&lt;STRONG&gt;Best Practice 6 — Review Cycles and Continuous Assessment&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;How do you know it's ready for users? We treat the application as a software product.&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;The Review Process:&lt;/STRONG&gt;&lt;SPAN&gt; For stakeholder feedback we review the app's captured conversations in MLflow Traces, where reviewers can inspect real interactions for tone, accuracy, and correctness before the production deployment. Note that the Databricks Review App Chat UI is built around a Model Serving endpoint, so it does not directly exercise an agent deployed as a Databricks App, if you want the interactive Review App experience for staging, deploy the agent to a separate Model Serving test endpoint and point the Review App at that.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Leveraging Assessments:&lt;/STRONG&gt;&lt;SPAN&gt; Don't just rely on manual testing. We use MLflow Agent Evaluation to run our "eval set" (a suite of 10+ common CPG business questions). By setting up a recurring job that runs these evaluations against new agent versions, we can catch regressions in correctness or safety before they hit production. Treat these "Judge" scores as probabilistic quality signals — a useful regression check, but a complement to, not a replacement for, deterministic unit tests, security tests, and integration tests. LLM judges are themselves non-deterministic and depend on the datasets, scorers, and thresholds you define.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;H3&gt;&lt;SPAN&gt;MLflow tracing &amp;amp; evaluation&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;mlflow.langchain.autolog()&lt;/SPAN&gt;&lt;SPAN&gt; captures the entire agent run. The trick for clean traces: wrap each invocation in a &lt;/SPAN&gt;&lt;STRONG&gt;parent span&lt;/STRONG&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt; (&lt;/SPAN&gt;&lt;SPAN&gt;&lt;a href="https://community.databricks.com/t5/user/viewprofilepage/user-id/171236"&gt;@mlflow&lt;/a&gt;.trace(name=..., span_type="AGENT")&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;&lt;FONT face="courier new,courier"&gt;)&lt;/FONT&gt;, otherwise autolog emits fragmented one-span-per-call traces. With the parent span you get a single unified trace:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;northstar_brand_copilot → LangGraph → ChatDatabricks → [Genie query OR Vector Search retrieve] → ChatDatabricks&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;For quality, &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;mlflow.genai.evaluate(...)&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; runs built-in judge scorers over a curated 10-question CPG eval set:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Safety 1.0 · RelevanceToQuery 0.9 · Correctness 0.8&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;Gotcha:&lt;/STRONG&gt;&lt;SPAN&gt; the GenAI judge scorers require the &lt;/SPAN&gt;&lt;SPAN&gt;databricks-agents&lt;/SPAN&gt;&lt;SPAN&gt; package. Without it, scorers fail &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;silently&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; — assessments come back &lt;/SPAN&gt;&lt;SPAN&gt;None&lt;/SPAN&gt;&lt;SPAN&gt;, which looks like passing tests but isn't. Install &lt;/SPAN&gt;&lt;SPAN&gt;databricks-agents&lt;/SPAN&gt;&lt;SPAN&gt; before evaluating.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;U&gt;&lt;STRONG&gt;MLFlow Trace:&lt;/STRONG&gt;&lt;/U&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_1-1785652299587.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/29671iBD1DC9DAA6A60567/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_1-1785652299587.png" alt="SashankKotta_1-1785652299587.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;U&gt;&lt;STRONG&gt;MLFlow Evaluations:&lt;/STRONG&gt;&lt;/U&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_0-1785666419071.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/29673iBD7513DFF4A242FD/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_0-1785666419071.png" alt="SashankKotta_0-1785666419071.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Deployment: Automation Bundles, step by step&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Everything is declared in &lt;/SPAN&gt;&lt;SPAN&gt;databricks.yml&lt;/SPAN&gt;&lt;SPAN&gt;. The app, its run command, every environment variable, and — crucially — &lt;/SPAN&gt;&lt;STRONG&gt;every resource grant&lt;/STRONG&gt;&lt;SPAN&gt; the service principal needs:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="markup"&gt;resources:
  apps:
    agent_langgraph:
      name: "northstar-brand-copilot"
      source_code_path: ./
      config:
        command: ["uv", "run", "start-server"]
        env:
          - { name: MODEL_ENDPOINT, value: "databricks-claude-sonnet-4-5" }
          - { name: GENIE_SPACE_ID, value: "01f1..." }
          - { name: LAKEBASE_INSTANCE_NAME, value_from: "database" }
      resources:
        - name: llm          # CAN_QUERY on the Claude endpoint
        - name: embedding    # CAN_QUERY on the embedding endpoint
        - name: genie_space  # CAN_RUN on the Genie space
        - name: vector_index # SELECT on the VS index (uc_securable)
        - name: database     # CAN_CONNECT_AND_CREATE on Lakebase
        - name: warehouse    # CAN_USE on the SQL warehouse
        - name: experiment   # CAN_MANAGE on the MLflow experiment&lt;/LI-CODE&gt;
&lt;P&gt;Note that running the app (&lt;FONT face="courier new,courier"&gt;databricks bundle run&lt;/FONT&gt;) after deployment is not optional. deploy (&lt;FONT face="courier new,courier"&gt;databricks bundle deploy&lt;/FONT&gt;) uploads code and reconciles resources; bundle run is what actually restarts the app, Avoiding it would be testing stale code.&lt;/P&gt;
&lt;H3&gt;&lt;SPAN&gt;The grants, bundle &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;can't&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; express&lt;/SPAN&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;While the bundle grants the SP access to endpoints like the Genie space, Vector Search index, and Lakebase instance, additional permissions are needed for execution. Specifically, since Genie runs SQL queries under the &lt;STRONG&gt;service principal's identity&lt;/STRONG&gt;, the SP must be granted Unity Catalog table access and SQL Warehouse usage. Similarly, Lakebase requires configuring a dedicated Postgres role. Those are applied by &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;STRONG&gt;deployment/grant_resources.py&lt;/STRONG&gt;&lt;/FONT&gt;&lt;SPAN&gt;:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;A.&lt;/STRONG&gt; &lt;SPAN&gt;USAGE&lt;/SPAN&gt;&lt;SPAN&gt; on the catalog/schema + &lt;/SPAN&gt;&lt;SPAN&gt;SELECT&lt;/SPAN&gt;&lt;SPAN&gt; on the tables (one statement per SQL Statement API call, it only runs one at a time)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;B.&lt;/STRONG&gt; &lt;SPAN&gt;CAN_USE&lt;/SPAN&gt;&lt;SPAN&gt; on the SQL warehouse&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;C.&lt;/STRONG&gt;&lt;SPAN&gt; A Lakebase Postgres role for the SP + grants on the &lt;/SPAN&gt;&lt;SPAN&gt;AsyncDatabricksStore&lt;/SPAN&gt;&lt;SPAN&gt; memory tables (run as a serverless job, because it needs &lt;/SPAN&gt;&lt;SPAN&gt;databricks_ai_bridge&lt;/SPAN&gt;&lt;SPAN&gt;)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;STRONG&gt;More gotchas from the trenches:&lt;/STRONG&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI&gt;
&lt;P data-path-to-node="3,1,0,0"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="3,1,0,0"&gt;Authentication:&lt;/STRONG&gt; Deployed Databricks Apps require an &lt;STRONG data-index-in-node="52" data-path-to-node="3,1,0,0"&gt;OAuth token&lt;/STRONG&gt; for API requests; Personal Access Tokens (PATs) will not work.&lt;/P&gt;
&lt;UL data-path-to-node="3,1,0,1"&gt;
&lt;LI&gt;
&lt;P data-path-to-node="3,1,0,1,0,0"&gt;&lt;I data-index-in-node="0" data-path-to-node="3,1,0,1,0,0"&gt;Quick fix:&lt;/I&gt; Retrieve your token using: &lt;FONT face="courier new,courier"&gt;databricks auth token ... | jq -r .access_token&lt;/FONT&gt;&lt;/P&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P data-path-to-node="3,1,1,0"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="3,1,1,0"&gt;YAML Configuration Syntax:&lt;/STRONG&gt; Under the &lt;FONT face="courier new,courier"&gt;sql_warehouse&lt;/FONT&gt; resource configuration, the correct property name is &lt;FONT face="courier new,courier"&gt;id:&lt;/FONT&gt;, not &lt;FONT face="courier new,courier"&gt;sql_warehouse_id:&lt;/FONT&gt; (a common typo in some documentation templates).&lt;/P&gt;
&lt;/LI&gt;
&lt;LI&gt;
&lt;P data-path-to-node="3,1,2,0"&gt;&lt;STRONG data-index-in-node="0" data-path-to-node="3,1,2,0"&gt;Terraform State Loss Recovery:&lt;/STRONG&gt; If your Terraform state is lost, running &lt;FONT face="courier new,courier"&gt;databricks bundle deploy&lt;/FONT&gt; will attempt to re-grant permissions on all resources, which fails on the Lakebase database grant.&lt;/P&gt;
&lt;UL data-path-to-node="3,1,2,1"&gt;
&lt;LI&gt;
&lt;P data-path-to-node="3,1,2,1,0,0"&gt;&lt;I data-index-in-node="0" data-path-to-node="3,1,2,1,0,0"&gt;Workaround:&lt;/I&gt; Because &lt;FONT face="courier new,courier"&gt;bundle deploy&lt;/FONT&gt; uploads the source code &lt;I data-index-in-node="58" data-path-to-node="3,1,2,1,0,0"&gt;before&lt;/I&gt; it fails during grant reconciliation, you can complete the deployment by running &lt;FONT face="courier new,courier"&gt;databricks apps deploy &amp;lt;app&amp;gt; --source-code-path &amp;lt;bundle files path&amp;gt;&lt;/FONT&gt; to update the application code directly without triggering resource reconciliation.&lt;/P&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;/LI&gt;
&lt;/UL&gt;
&lt;H2&gt;&lt;SPAN&gt;Final App:&lt;/SPAN&gt;&lt;/H2&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;U&gt;Figure 1(Dashbaord):&lt;/U&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_3-1783743763523.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/28860iC550670D973FB85B/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_3-1783743763523.png" alt="SashankKotta_3-1783743763523.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;U&gt;&lt;STRONG&gt;Figure 2(Agent):&lt;/STRONG&gt;&lt;/U&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_4-1783743763524.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/28859iD685EB10E2F7C423/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_4-1783743763524.png" alt="SashankKotta_4-1783743763524.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;U&gt;&lt;STRONG&gt;Figure 3(Agent Memory: Saved to Lakebase):&lt;/STRONG&gt;&lt;/U&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_5-1783743763524.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/28861i5455357D2793F8F8/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_5-1783743763524.png" alt="SashankKotta_5-1783743763524.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Performance Considerations — Latency&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;Latency is a critical metric for a responsive Copilot experience. Since the agent invokes multiple tools and LLM endpoints sequentially, we prioritize streaming responses. By utilizing streaming responses and optimizing tool execution, we ensure that the initial token latency is kept to a minimum, providing a snappy, responsive interface.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;SPAN&gt;Lessons learned&lt;/SPAN&gt;&lt;/H2&gt;
&lt;OL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Prefer the App + ResponsesAgent pattern&lt;/STRONG&gt;&lt;SPAN&gt; over log-to-UC + Model Serving for new agents, It unifies UI, API, identity, and observability.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Your system prompt and tool descriptions are your routing engine.&lt;/STRONG&gt;&lt;SPAN&gt; Spend time there before writing graph code.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Consume Databricks services as MCP tools&lt;/STRONG&gt;&lt;SPAN&gt; — managed endpoints, governed identity, zero glue code.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Normalize tool output&lt;/STRONG&gt;&lt;SPAN&gt; to your LLM provider's expectations (the Genie &lt;/SPAN&gt;&lt;SPAN&gt;id&lt;/SPAN&gt;&lt;SPAN&gt;-field fix).&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Make every external dependency degrade gracefully&lt;/STRONG&gt;&lt;SPAN&gt; — tools and memory both.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Treat memory as a first-class, per-user, namespaced store&lt;/STRONG&gt;&lt;SPAN&gt; with explicit &lt;/SPAN&gt;&lt;SPAN&gt;user_id&lt;/SPAN&gt;&lt;SPAN&gt; handling and JSON validation.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Declare all grants in the bundle, and know which ones it can't express&lt;/STRONG&gt;&lt;SPAN&gt; (UC tables, warehouse, Lakebase role).&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Wrap invocations in a parent MLflow span&lt;/STRONG&gt;&lt;SPAN&gt; and install &lt;/SPAN&gt;&lt;SPAN&gt;databricks-agents&lt;/SPAN&gt;&lt;SPAN&gt; for real evaluation scores.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;SPAN&gt;The result is a governed, observable, reproducible agent that answers real CPG business questions — and a blueprint you can lift for your own domain.&lt;/SPAN&gt;&lt;/P&gt;</description>
    <pubDate>Mon, 10 Aug 2026 11:43:09 GMT</pubDate>
    <dc:creator>SashankKotta</dc:creator>
    <dc:date>2026-08-10T11:43:09Z</dc:date>
    <item>
      <title>Building a Production LangGraph Agent on Databricks - NorthStar Brand Copilot</title>
      <link>https://community.databricks.com/t5/technical-blog/building-a-production-langgraph-agent-on-databricks-northstar/ba-p/162634</link>
      <description>&lt;P&gt;&lt;SPAN&gt;We built &lt;/SPAN&gt;&lt;STRONG&gt;NorthStar Brand Copilot&lt;/STRONG&gt;&lt;SPAN&gt;, an AI assistant for brand managers and field-sales reps at a fictional multi-category CPG company (Snacks, Beverages, Personal Care). It's a &lt;/SPAN&gt;&lt;STRONG&gt;LangGraph custom agent&lt;/STRONG&gt;&lt;SPAN&gt; that routes each question to the right &lt;/SPAN&gt;&lt;STRONG&gt;Databricks-native&lt;/STRONG&gt;&lt;SPAN&gt; capability:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Genie&lt;/STRONG&gt;&lt;SPAN&gt; (NatualLanguage→SQL) for the numbers — sell-in/sell-out, trade-promotion ROI, inventory, market share&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;AI Search&lt;/STRONG&gt;(&lt;SPAN&gt;RAG) for the documents — product specs, allergens, consumer reviews, brand guidelines, the promo playbook&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Lakebase&lt;/STRONG&gt;&lt;SPAN&gt; (long-term memory) for the decisions, For ex: "remember we cut BOGO (Buy One Get One) at Walgreens," recalled across sessions. This provides a persistent, semantic context layer, ensuring the agent doesn't process queries in isolation but evolves its understanding based on past decisions and specific user history.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;LangGraph&lt;/STRONG&gt;&lt;SPAN&gt; for agent orchestration. While simpler frameworks (like standard LangChain chains or monolithic agent patterns) are easier to start with, LangGraph was chosen for its superior ability to handle cyclical, complex workflows and granular state management. It allows the agent to reason, refine its tool-calling strategy iteratively, and maintain persistent state, which is essential for enterprise-grade reliability compared to more rigid, linear alternatives.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;&lt;STRONG&gt;Note&lt;/STRONG&gt;: The CPG organization and data used in this blog post are fictional and intended for&amp;nbsp;&lt;/SPAN&gt;&lt;/I&gt;&lt;I&gt;&lt;SPAN&gt;demonstration purposes only.&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="SashankKotta_0-1785652864446.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/29672i2445C22B1BE89A55/image-size/large?v=v2&amp;amp;px=999" role="button" title="SashankKotta_0-1785652864446.png" alt="SashankKotta_0-1785652864446.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&amp;nbsp;&lt;/P&gt;</description>
      <pubDate>Mon, 10 Aug 2026 11:43:09 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/building-a-production-langgraph-agent-on-databricks-northstar/ba-p/162634</guid>
      <dc:creator>SashankKotta</dc:creator>
      <dc:date>2026-08-10T11:43:09Z</dc:date>
    </item>
    <item>
      <title>Re: Building a Production LangGraph Agent on Databricks - NorthStar Brand Copilot</title>
      <link>https://community.databricks.com/t5/technical-blog/building-a-production-langgraph-agent-on-databricks-northstar/bc-p/165511#M1170</link>
      <description>&lt;P&gt;Great end-to-end write-up — the MLflow tracing setup in particular is a pattern more teams should adopt from day one rather than bolting on later.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;One thing worth extending that tracing to: per-node cost attribution. Since every ChatDatabricks call in the LangGraph graph is a separate model invocation, the context window grows with each tool result appended to the message list. In a session where the agent calls Genie + Vector Search + Lakebase in sequence, you're often paying 3–5× what a single-query baseline would suggest — because each tool output accumulates in the prompt for all subsequent steps.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;MLflow spans already capture token counts per step. Adding a cost_usd attribute to each span (input_tokens * price_in + output_tokens * price_out) and grouping by session_id gives you a per-conversation cost view that's essential for tuning buffer sizes before you hit production scale.&lt;/P&gt;&lt;P&gt;&amp;nbsp;&lt;/P&gt;&lt;P&gt;Two optimizations that help: (1) Summarize each tool response before appending it to the graph state — passing a 50-token summary instead of a 500-token raw Genie result keeps context lean across subsequent nodes. (2) Not all nodes need Sonnet-class reasoning. The routing step (which tool to call?) is usually a simpler classification task than the final answer synthesis. Splitting that to a lighter model and reserving claude-sonnet-4-5 for the final generation step can cut per-session cost 30–50% with no visible quality regression on the routing decision itself.&lt;/P&gt;</description>
      <pubDate>Wed, 12 Aug 2026 14:38:43 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/building-a-production-langgraph-agent-on-databricks-northstar/bc-p/165511#M1170</guid>
      <dc:creator>DoTA</dc:creator>
      <dc:date>2026-08-12T14:38:43Z</dc:date>
    </item>
  </channel>
</rss>

