<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>article Combining Agent Bricks and ai_parse_document for Smarter Document Workflows in Technical Blog</title>
    <link>https://community.databricks.com/t5/technical-blog/combining-agent-bricks-and-ai-parse-document-for-smarter/ba-p/143801</link>
    <description>&lt;H1&gt;Building a Research Paper Curator for Knowledge Assistants&lt;/H1&gt;
&lt;P&gt;&lt;SPAN&gt;Document-heavy workflows such as research analysis, contract review, and support ticket processing, are a natural fit for AI agents. Building them from scratch means orchestrating document parsing, information extraction, vector search, chat interfaces, and likely much more. That’s a lot of plumbing before you get to the interesting part.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Databricks Agent Bricks and the &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; SQL function eliminate most of that plumbing. In this post, we’ll show how these tools work together by building an application for curating research papers for a knowledge assistant. Using this example, we’ll highlight patterns you can apply to your own agentic applications on Databricks.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;What we’ll use:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;The &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; Databricks SQL function for extracting structured text from PDFs&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/generative-ai/agent-bricks/key-info-extraction" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Agent Bricks Key Information Extraction&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; for pulling structured fields from documents&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/generative-ai/agent-bricks/knowledge-assistant" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Agent Bricks Knowledge Assistant&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; for RAG-based Q&amp;amp;A&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;A href="https://docs.databricks.com/aws/en/dev-tools/databricks-apps/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Databricks Apps&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; for the application interface&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;This is a high-level structural overview, so we won’t walk through every line of code. Instead, we’ll focus on how these components interact and unlock value together. For the full implementation, &lt;/SPAN&gt;&lt;A href="https://github.com/databricks-solutions/devrel-examples/tree/main/demo_projects/arxiv" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;check out the project on GitHub&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;The Problem: Knowledge Assistants Need Curation&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Knowledge assistants such as the &lt;/SPAN&gt;&lt;A href="https://docs.databricks.com/aws/en/generative-ai/agent-bricks/knowledge-assistant" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Databricks Agent Bricks Knowledge Assistant&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; work best when they have access to focused, relevant content. Overloading them with irrelevant materials can lead to knowledge assistants retrieving only some of the relevant information they need to address a query or, worse, retrieving entirely incorrect information (e.g., &lt;/SPAN&gt;&lt;A href="https://aclanthology.org/2025.acl-long.131/" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;1&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;A href="https://arxiv.org/abs/2601.01896" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;2&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;).&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Manually pre-screening materials for inclusion in a knowledge assistant can be a cumbersome and time-consuming task, especially when it comes to technical documents like research papers. You may need to search through dozens of pages of text to determine whether a paper warrants inclusion in your knowledge assistant.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;The Solution: A Curation Workflow&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;In order to make the screening process as painless as possible, we can make use of several AI features on Databricks to pull out some key details from the research papers and a Databricks app to create a simple UI for flagging papers to include.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_1-1768248344551.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22912i4B027061FFCE537A/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_1-1768248344551.png" alt="DanielLiden_1-1768248344551.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;The curator app: search arXiv, parse and extract with AI, manage your knowledge base&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;To accomplish this, we will:&lt;/SPAN&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Search the arXiv API for papers based on topics and keywords&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Get the text from promising candidates with the &lt;/SPAN&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;SPAN&gt; SQL function&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Extract key details from the candidate papers with Agent Bricks Information Extraction&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Review the extracted fields, which are targeted and relevant to the criteria we’re using to filter papers for inclusion in our Knowledge Assistant. This is much quicker than manual review!&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Add relevant papers to the knowledge assistant.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_0-1768248344550.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22911i5FBDEA9573BA4244/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_0-1768248344550.png" alt="DanielLiden_0-1768248344550.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;EM&gt;Document search, parsing, and review workflow for knowledge assistant curation&lt;/EM&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Setting Up the Agents&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;Agent Bricks agents are tailored to specific tasks and data, so we need to initialize the Information Extraction and Knowledge Assistant agents before we can integrate them into our application.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;After the agents are configured (via the Databricks UI), we will invoke them via the OpenAI SDK. Agent Bricks endpoints are OpenAI-compatible, so you can use the familiar OpenAI SDK pattern:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="python"&gt;from databricks.sdk import WorkspaceClient

ws_client = WorkspaceClient()
openai_client = ws_client.serving_endpoints.get_open_ai_client()

response = openai_client.chat.completions.create(
    model="your-agent-endpoint",
    messages=[{"role": "user", "content": "Your prompt here"}],
)&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;This pattern works for both KIE and Knowledge Assistant endpoints. The &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;WorkspaceClient&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; handles authentication automatically.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Here’s how to set up the Agent Bricks agents.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Key Information Extraction (KIE)&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;The KIE agent extracts structured fields from parsed documents, giving you a quick preview of each paper’s contributions, methodology, and limitations without reading 30 pages. You configure what to extract by defining a JSON schema, then point the agent at your parsed text. The &lt;/SPAN&gt;&lt;A href="https://docs.databricks.com/aws/en/generative-ai/agent-bricks/key-info-extraction" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;full KIE documentation&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; covers all configuration options.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Preparing the data. Before KIE can extract anything, PDFs need to be converted to text. We started with a golden set of seminal LLM agent papers (ReAct, Reflexion, etc.) and parsed them using the &lt;/SPAN&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;SPAN&gt; SQL function:&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;STRONG&gt;&lt;FONT face="courier new,courier"&gt;SELECT ai_parse_document(content) as parsed&lt;/FONT&gt;&lt;/STRONG&gt;&lt;BR /&gt;&lt;STRONG&gt;&lt;FONT face="courier new,courier"&gt;FROM read_files('/Volumes/catalog/schema/volume/paper.pdf')&lt;/FONT&gt;&lt;/STRONG&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;The &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; function handles multi-column layouts, equations, and citations, returning structured JSON with page-by-page content. We stored the extracted text in a &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;parsed_documents&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; table, which becomes the KIE agent’s data source.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Defining the schema. The Agent Bricks UI generates a starter schema based on your data, but you’ll want to replace it with fields relevant to your use case. Switch to JSON schema mode and define what you need. For paper screening, we defined seven fields: &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;title&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;authors&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;affiliation&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;methodology&lt;/FONT&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;contributions&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;, &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;limitations&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;, and &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;topics&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;. Here’s an excerpt:&lt;/SPAN&gt;&lt;/P&gt;
&lt;LI-CODE lang="javascript"&gt;{
  "properties": {
    ...
    "methodology": {
      "type": "string",
      "description": "Research methods, model architectures, and experimental approaches used."
    },
    "limitations": {
      "type": "array",
      "items": {"type": "string"},
      "description": "Acknowledged weaknesses, constraints, and areas for improvement."
    },
    ...
  }
}
&lt;/LI-CODE&gt;
&lt;P&gt;&lt;SPAN&gt;The &lt;/SPAN&gt;&lt;SPAN&gt;description&lt;/SPAN&gt;&lt;SPAN&gt; fields matter: they guide the model’s extraction logic. If you’re getting inconsistent results, refining the description often helps more than adding examples. We also used &lt;/SPAN&gt;&lt;SPAN&gt;anyOf&lt;/SPAN&gt;&lt;SPAN&gt; with null for optional fields (like &lt;/SPAN&gt;&lt;SPAN&gt;affiliation&lt;/SPAN&gt;&lt;SPAN&gt;) so the model returns &lt;/SPAN&gt;&lt;SPAN&gt;null&lt;/SPAN&gt;&lt;SPAN&gt; when information isn’t present rather than hallucinating.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_2-1768248344552.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22913iF7A4F5C4502EB50F/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_2-1768248344552.png" alt="DanielLiden_2-1768248344552.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;KIE agent configuration showing extraction schema and sample outputs&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Knowledge Assistant&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;The Knowledge Assistant provides RAG-based Q&amp;amp;A over your documents. Setup is straightforward: point it at a Unity Catalog Volume and deploy. The &lt;/SPAN&gt;&lt;A href="https://docs.databricks.com/aws/en/generative-ai/agent-bricks/knowledge-assistant" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;full Knowledge Assistant documentation&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt; covers additional options like vector search indexes and custom instructions.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;Configuring the data source. In the Agent Bricks UI, select Knowledge Assistant and click Build. You’ll need to specify:&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Name: We used &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;arxiv-papers&lt;/SPAN&gt;&lt;/FONT&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Knowledge Source: Select Unity Catalog Volume and point it at your PDFs volume (e.g., &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;arxiv_demo.main.pdfs&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt;)&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Instructions (optional): Guidelines for how the agent should respond&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;Only documents in this volume will be indexed. For the curator application, we set up a separate staging volume for papers under review. documents only move to the KA volume when explicitly approved. This keeps the knowledge base focused.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;After deployment, the agent takes a few minutes to sync and index your documents. You can test it immediately via the embedded chat or AI Playground. If you add or remove files from the volume later, click Sync in the UI to update the index.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_3-1768248344553.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22915i6D91346B47B7B480/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_3-1768248344553.png" alt="DanielLiden_3-1768248344553.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;Knowledge Assistant configured with a UC Volume source&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;The Curation Workflow&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;We built a Databricks App with a four-phase workflow that calls upon these agents:&lt;/SPAN&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Phase 1: Search&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;Search arXiv for papers matching your criteria. The UI lets you filter by category, date range, and keywords. You can select papers that look like they &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;might&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; be relevant to your interests and worth including in the knowledge assistant.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_4-1768248344553.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22914i2B02CB70DF688C79/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_4-1768248344553.png" alt="DanielLiden_4-1768248344553.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;Search interface with category filters and date range&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Phase 2: Review (Parse + Extract)&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;Once you select some papers, you can click the “Parse Papers” button beneath the list. This triggers the following:&lt;/SPAN&gt;&lt;/P&gt;
&lt;OL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;The selected papers are downloaded and uploaded to a staging volume. Papers in the staging volume are &lt;/SPAN&gt;&lt;I&gt;&lt;SPAN&gt;not&lt;/SPAN&gt;&lt;/I&gt;&lt;SPAN&gt; accessible by the Knowledge Assistant.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;The &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; SQL function is used to obtain the text of the papers and save them to a Unity Catalog table.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;The Information Extraction agent we defined previously extracts the relevant fields from the papers.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/OL&gt;
&lt;P&gt;&lt;SPAN&gt;Once the parsing and extraction are complete (which might take a few minutes, depending on the number of papers), you can review the key contributions, methods, and limitations of each paper in the review tab. From the review tab, you can select which papers to add to the knowledge assistant.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_5-1768248344554.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22916iFBFBAC8859443706/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_5-1768248344554.png" alt="DanielLiden_5-1768248344554.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;KIE-extracted insights showing key contributions and methodology&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Phase 3: Manage&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;The KA Manager shows what’s currently in your knowledge base. You can add or remove documents. Keeping the knowledge assistant up to date with relevant materials helps to keep the results targeted and relevant over time.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_6-1768248344555.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22918iB0A80E28661F93B0/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_6-1768248344555.png" alt="DanielLiden_6-1768248344555.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;Knowledge Assistant Manager showing curated papers&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;H3&gt;&lt;STRONG&gt;Phase 4: Chat&lt;/STRONG&gt;&lt;/H3&gt;
&lt;P&gt;&lt;SPAN&gt;Query your curated knowledge base. Because you’ve been selective about what goes in, responses are focused and relevant.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;span class="lia-inline-image-display-wrapper lia-image-align-inline" image-alt="DanielLiden_7-1768248344555.png" style="width: 999px;"&gt;&lt;img src="https://community.databricks.com/t5/image/serverpage/image-id/22917i0C766DE5808E14E1/image-size/large?v=v2&amp;amp;px=999" role="button" title="DanielLiden_7-1768248344555.png" alt="DanielLiden_7-1768248344555.png" /&gt;&lt;/span&gt;&lt;/P&gt;
&lt;P&gt;&lt;I&gt;&lt;SPAN&gt;Chat interface querying the curated knowledge base&lt;/SPAN&gt;&lt;/I&gt;&lt;/P&gt;
&lt;P&gt;&lt;SPAN&gt;There are several other ways you can chat with the knowledge assistant. You can use the Knowledge Assistant UI directly in Databricks by clicking on the “Agents” tab, selecting your knowledge assistant, and using the Knowledge Assistant chat interface. You can also use the Databricks AI Playground by clicking the Playground tab and finding your Knowledge Assistant endpoint in the model selection dropdown. We included the chat interface directly in the curator application primarily for convenience.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Key Takeaway&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;SPAN&gt;We used a paper curator as our example, but the real takeaway is how quickly you can build document-processing agents on Databricks. Agent Bricks + &lt;/SPAN&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; handle the hard parts, so you can focus on your use case.&lt;/SPAN&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Knowledge Assistant gives you a Q&amp;amp;A bot over your documents in just a few UI-driven steps, without needing to worry about configuring a vector database or chat interface yourself.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;The Information Extraction agent can obtain key details from complex documents and return them in a structure you define. Again, this takes only a few UI-driven steps.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;FONT face="courier new,courier"&gt;&lt;SPAN&gt;ai_parse_document&lt;/SPAN&gt;&lt;/FONT&gt;&lt;SPAN&gt; makes it easy to extract clean, structured text from documents and excels at complicated documents like research papers that may include tables and figures.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;SPAN&gt;Databricks Apps let you define custom application logic for tying all the pieces together.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;In this example, we developed a curator app that simplifies the process of adding relevant data to a knowledge assistant. But the broader principles apply across use cases. With Databricks Agent Bricks, creating performant, high-quality agents that integrate seamlessly with your applications and workflows has never been easier, and they’re even more powerful when used together.&lt;/SPAN&gt;&lt;/P&gt;
&lt;H2&gt;&lt;STRONG&gt;Try It Yourself&lt;/STRONG&gt;&lt;/H2&gt;
&lt;P&gt;&lt;A href="https://github.com/databricks-solutions/devrel-examples/tree/main/demo_projects/arxiv" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;Check out the project on GitHub&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;. To get started, import the repository to your Databricks workspace and follow the steps in the project’s &lt;/SPAN&gt;&lt;A href="https://github.com/databricks-solutions/devrel-examples/blob/main/demo_projects/arxiv/Runbook.ipynb" target="_blank" rel="noopener"&gt;&lt;SPAN&gt;runbook&lt;/SPAN&gt;&lt;/A&gt;&lt;SPAN&gt;. This will guide you through the process of setting up the agents, creating your agents, and deploying the application.&lt;/SPAN&gt;&lt;/P&gt;</description>
    <pubDate>Mon, 12 Jan 2026 20:58:11 GMT</pubDate>
    <dc:creator>Daniel-Liden</dc:creator>
    <dc:date>2026-01-12T20:58:11Z</dc:date>
    <item>
      <title>Combining Agent Bricks and ai_parse_document for Smarter Document Workflows</title>
      <link>https://community.databricks.com/t5/technical-blog/combining-agent-bricks-and-ai-parse-document-for-smarter/ba-p/143801</link>
      <description>&lt;P&gt;Document workflows need a lot of plumbing, such as parsing, extraction, vector search, and chat interfaces, before you get meaningful and usable results. Agent Bricks and ai_parse_document handle the hard parts. We built a research paper curator to show how these tools work together, but the patterns apply wherever you're turning documents into decisions.&lt;/P&gt;</description>
      <pubDate>Mon, 12 Jan 2026 20:58:11 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/combining-agent-bricks-and-ai-parse-document-for-smarter/ba-p/143801</guid>
      <dc:creator>Daniel-Liden</dc:creator>
      <dc:date>2026-01-12T20:58:11Z</dc:date>
    </item>
    <item>
      <title>Re: Combining Agent Bricks and ai_parse_document for Smarter Document Workflows</title>
      <link>https://community.databricks.com/t5/technical-blog/combining-agent-bricks-and-ai-parse-document-for-smarter/bc-p/144100#M884</link>
      <description>&lt;P&gt;this is fire. thanks for sharing&lt;/P&gt;</description>
      <pubDate>Wed, 14 Jan 2026 22:19:33 GMT</pubDate>
      <guid>https://community.databricks.com/t5/technical-blog/combining-agent-bricks-and-ai-parse-document-for-smarter/bc-p/144100#M884</guid>
      <dc:creator>npitts</dc:creator>
      <dc:date>2026-01-14T22:19:33Z</dc:date>
    </item>
  </channel>
</rss>

