cancel
Showing results for 
Search instead for 
Did you mean: 
Announcements
Stay up-to-date with the latest announcements from Databricks. Learn about product updates, new features, and important news that impact your data analytics workflow.
cancel
Showing results for 
Search instead for 
Did you mean: 

Announcement | Introducing OfficeQA Pro V2: Grounded-Reasoning Benchmark for Enterprise Data

Tushar_Parekar
Databricks Employee
Databricks Employee

Databricks has released OfficeQA Pro V2, a new benchmark for testing how well AI agents answer complex questions using evidence from large, unfamiliar document collections. Built from more than 1,400 U.S. Treasury PDFs spanning 1793–2024, the benchmark is designed to measure grounded reasoning in enterprise-style tasks, not just general question answering.

Key highlights

  • A demanding new benchmark: OfficeQA Pro V2 includes 90 questions over roughly 120,000 pages, with a median of 5.5 source documents per question. Many questions require retrieval, calculations, changing reporting conventions, and careful use of evidence.
  • Built to test generalization: The benchmark uses a new corpus so teams can evaluate whether improvements carry over to unfamiliar documents and tasks, rather than optimizing for one known dataset.
  • Created with synthetic data and verification: Databricks used a synthetic-data pipeline with automated checks, independent solver agents, and human review to create questions with traceable answers.
  • Agent harnesses matter: In Databricks’ evaluation, out-of-the-box agents averaged 26% accuracy, while Genie improved results by an average of 24 percentage points across matched model comparisons, reaching up to 60% accuracy in the strongest configuration.
  • Still a hard problem: The results show meaningful progress, but also continued challenges with document parsing, retrieval, historical revisions, and analytical reasoning.

OfficeQA Pro V2 is now available on Hugging Face, with evaluation code available on GitHub. Teams can use it to compare models and agent architectures, identify where failures occur, and measure how changes affect accuracy, cost, and latency before applying them to their own enterprise data.

👉 Read the full post here

0 REPLIES 0