<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:taxo="http://purl.org/rss/1.0/modules/taxonomy/" version="2.0">
  <channel>
    <title>topic Announcement | Introducing OfficeQA Pro V2: Grounded-Reasoning Benchmark for Enterprise Data in Announcements</title>
    <link>https://community.databricks.com/t5/announcements/announcement-introducing-officeqa-pro-v2-grounded-reasoning/m-p/165045#M981</link>
    <description>&lt;P&gt;&lt;SPAN&gt;Databricks has released &lt;/SPAN&gt;&lt;STRONG&gt;OfficeQA Pro V2&lt;/STRONG&gt;&lt;SPAN&gt;, a new benchmark for testing how well AI agents answer complex questions using evidence from large, unfamiliar document collections. Built from more than &lt;/SPAN&gt;&lt;STRONG&gt;1,400 U.S. Treasury PDFs&lt;/STRONG&gt;&lt;SPAN&gt; spanning &lt;/SPAN&gt;&lt;STRONG&gt;1793–2024&lt;/STRONG&gt;&lt;SPAN&gt;, the benchmark is designed to measure grounded reasoning in enterprise-style tasks, not just general question answering.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;Key highlights&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;A demanding new benchmark&lt;/STRONG&gt;&lt;SPAN&gt;: OfficeQA Pro V2 includes &lt;/SPAN&gt;&lt;STRONG&gt;90 questions&lt;/STRONG&gt;&lt;SPAN&gt; over roughly &lt;/SPAN&gt;&lt;STRONG&gt;120,000 pages&lt;/STRONG&gt;&lt;SPAN&gt;, with a median of &lt;/SPAN&gt;&lt;STRONG&gt;5.5 source documents per question&lt;/STRONG&gt;&lt;SPAN&gt;. Many questions require retrieval, calculations, changing reporting conventions, and careful use of evidence.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Built to test generalization&lt;/STRONG&gt;&lt;SPAN&gt;: The benchmark uses a new corpus so teams can evaluate whether improvements carry over to unfamiliar documents and tasks, rather than optimizing for one known dataset.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Created with synthetic data and verification&lt;/STRONG&gt;&lt;SPAN&gt;: Databricks used a synthetic-data pipeline with automated checks, independent solver agents, and human review to create questions with traceable answers.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Agent harnesses matter&lt;/STRONG&gt;&lt;SPAN&gt;: In Databricks’ evaluation, out-of-the-box agents averaged &lt;/SPAN&gt;&lt;STRONG&gt;26% accuracy&lt;/STRONG&gt;&lt;SPAN&gt;, while Genie improved results by an average of &lt;/SPAN&gt;&lt;STRONG&gt;24 percentage points&lt;/STRONG&gt;&lt;SPAN&gt; across matched model comparisons, reaching up to &lt;/SPAN&gt;&lt;STRONG&gt;60% accuracy&lt;/STRONG&gt;&lt;SPAN&gt; in the strongest configuration.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Still a hard problem&lt;/STRONG&gt;&lt;SPAN&gt;: The results show meaningful progress, but also continued challenges with document parsing, retrieval, historical revisions, and analytical reasoning.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;OfficeQA Pro V2 is now available on &lt;/SPAN&gt;&lt;STRONG&gt;Hugging Face&lt;/STRONG&gt;&lt;SPAN&gt;, with evaluation code available on GitHub. Teams can use it to compare models and agent architectures, identify where failures occur, and measure how changes affect accuracy, cost, and latency before applying them to their own enterprise data.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P class="p8i6j01 paragraph"&gt;&lt;A style="background-color: #ff3621; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px; font-weight: bold; display: inline-block;" href="https://www.databricks.com/blog/introducing-officeqa-pro-v2-new-benchmark-enterprise-grounded-reasoning?utm_source=bambu&amp;amp;utm_medium=social&amp;amp;utm_campaign=advocacy" target="_blank" rel="noopener"&gt; &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_right:"&gt;👉&lt;/span&gt; Read the full post here &lt;/A&gt;&lt;/P&gt;</description>
    <pubDate>Thu, 06 Aug 2026 17:14:36 GMT</pubDate>
    <dc:creator>Tushar_Parekar</dc:creator>
    <dc:date>2026-08-06T17:14:36Z</dc:date>
    <item>
      <title>Announcement | Introducing OfficeQA Pro V2: Grounded-Reasoning Benchmark for Enterprise Data</title>
      <link>https://community.databricks.com/t5/announcements/announcement-introducing-officeqa-pro-v2-grounded-reasoning/m-p/165045#M981</link>
      <description>&lt;P&gt;&lt;SPAN&gt;Databricks has released &lt;/SPAN&gt;&lt;STRONG&gt;OfficeQA Pro V2&lt;/STRONG&gt;&lt;SPAN&gt;, a new benchmark for testing how well AI agents answer complex questions using evidence from large, unfamiliar document collections. Built from more than &lt;/SPAN&gt;&lt;STRONG&gt;1,400 U.S. Treasury PDFs&lt;/STRONG&gt;&lt;SPAN&gt; spanning &lt;/SPAN&gt;&lt;STRONG&gt;1793–2024&lt;/STRONG&gt;&lt;SPAN&gt;, the benchmark is designed to measure grounded reasoning in enterprise-style tasks, not just general question answering.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P&gt;&lt;FONT size="4"&gt;&lt;STRONG&gt;Key highlights&lt;/STRONG&gt;&lt;/FONT&gt;&lt;/P&gt;
&lt;UL&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;A demanding new benchmark&lt;/STRONG&gt;&lt;SPAN&gt;: OfficeQA Pro V2 includes &lt;/SPAN&gt;&lt;STRONG&gt;90 questions&lt;/STRONG&gt;&lt;SPAN&gt; over roughly &lt;/SPAN&gt;&lt;STRONG&gt;120,000 pages&lt;/STRONG&gt;&lt;SPAN&gt;, with a median of &lt;/SPAN&gt;&lt;STRONG&gt;5.5 source documents per question&lt;/STRONG&gt;&lt;SPAN&gt;. Many questions require retrieval, calculations, changing reporting conventions, and careful use of evidence.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Built to test generalization&lt;/STRONG&gt;&lt;SPAN&gt;: The benchmark uses a new corpus so teams can evaluate whether improvements carry over to unfamiliar documents and tasks, rather than optimizing for one known dataset.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Created with synthetic data and verification&lt;/STRONG&gt;&lt;SPAN&gt;: Databricks used a synthetic-data pipeline with automated checks, independent solver agents, and human review to create questions with traceable answers.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Agent harnesses matter&lt;/STRONG&gt;&lt;SPAN&gt;: In Databricks’ evaluation, out-of-the-box agents averaged &lt;/SPAN&gt;&lt;STRONG&gt;26% accuracy&lt;/STRONG&gt;&lt;SPAN&gt;, while Genie improved results by an average of &lt;/SPAN&gt;&lt;STRONG&gt;24 percentage points&lt;/STRONG&gt;&lt;SPAN&gt; across matched model comparisons, reaching up to &lt;/SPAN&gt;&lt;STRONG&gt;60% accuracy&lt;/STRONG&gt;&lt;SPAN&gt; in the strongest configuration.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;LI style="font-weight: 400;" aria-level="1"&gt;&lt;STRONG&gt;Still a hard problem&lt;/STRONG&gt;&lt;SPAN&gt;: The results show meaningful progress, but also continued challenges with document parsing, retrieval, historical revisions, and analytical reasoning.&lt;/SPAN&gt;&lt;/LI&gt;
&lt;/UL&gt;
&lt;P&gt;&lt;SPAN&gt;OfficeQA Pro V2 is now available on &lt;/SPAN&gt;&lt;STRONG&gt;Hugging Face&lt;/STRONG&gt;&lt;SPAN&gt;, with evaluation code available on GitHub. Teams can use it to compare models and agent architectures, identify where failures occur, and measure how changes affect accuracy, cost, and latency before applying them to their own enterprise data.&lt;/SPAN&gt;&lt;/P&gt;
&lt;P class="p8i6j01 paragraph"&gt;&lt;A style="background-color: #ff3621; color: white; padding: 10px 20px; text-decoration: none; border-radius: 5px; font-weight: bold; display: inline-block;" href="https://www.databricks.com/blog/introducing-officeqa-pro-v2-new-benchmark-enterprise-grounded-reasoning?utm_source=bambu&amp;amp;utm_medium=social&amp;amp;utm_campaign=advocacy" target="_blank" rel="noopener"&gt; &lt;span class="lia-unicode-emoji" title=":backhand_index_pointing_right:"&gt;👉&lt;/span&gt; Read the full post here &lt;/A&gt;&lt;/P&gt;</description>
      <pubDate>Thu, 06 Aug 2026 17:14:36 GMT</pubDate>
      <guid>https://community.databricks.com/t5/announcements/announcement-introducing-officeqa-pro-v2-grounded-reasoning/m-p/165045#M981</guid>
      <dc:creator>Tushar_Parekar</dc:creator>
      <dc:date>2026-08-06T17:14:36Z</dc:date>
    </item>
  </channel>
</rss>

