OFICIAL Databricks Newsroom

Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning

What happened
Based on Databricks Newsroom · Aug 06, 2026

Databricks released OfficeQA Pro V2, a benchmark testing AI agents' ability to generalize grounded reasoning across unfamiliar enterprise documents, following the original OfficeQA benchmark introduced seven months prior.

Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning
Databricks Newsroom — Databricks
Key points
·
Today, the company is releasing OfficeQA Pro V2, a new benchmark designed to evaluate whether AI agents can generalize to unfamiliar, enterprise-style grounded-reasoning tasks.
·
Seven months ago, we introduced the OfficeQA benchmark to measure how well AI systems answer analytical questions using evidence from large document collections, an extremely common and important enterprise task that we found agents struggled with.
·
Since its introduction, OfficeQA, and its frontier subset, OfficeQA Pro, has become an important measure for frontier model and agent capabilities, driving progress in document retrieval, parsing, and analytical reasoning.
·
But this progress raises a fundamental question: do these improvements reflect broader advances in grounded reasoning, or progress specific to one corpus and task distribution?
Key numbers
·
1%, with the winning team reaching 63.
·
3%.
·
0 percentage points, with its strongest configuration reaching 60%.

OfficeQA Pro V2 evaluates whether AI systems can generalize grounded reasoning beyond a single document collection, addressing a gap identified in enterprise settings where agents rarely operate on stable corpora. The benchmark was first used in the Databricks Grounded Reasoning Cup, a competition involving 11 academic teams supported by OpenAI, Anthropic, and Google DeepMind, evaluated on a previously unseen corpus of approximately 120,000 pages from the U.S. Treasury’s Accounts of Receipts and Expenditures. The dataset, released by the U.S. Treasury for the first time to mark the 250th anniversary of the United States, spans accounting records from 1793 to 2024, including dense tables, evolving conventions, and institutional complexities typical of enterprise documents.

Out-of-the-box frontier agents achieved an average accuracy of 37.5% on OfficeQA Pro V2, while agents developed specifically for the competition averaged 41.1%, with the winning team reaching 63.3%. Databricks’ Genie agent, using the same underlying models, improved accuracy by an average of 24.0 percentage points, with its strongest configuration reaching 60%. Performance varied significantly by model and cost, with higher-priced models not consistently outperforming cheaper ones, highlighting the role of agent harness design in improving results.

OfficeQA Pro V2 introduces more rigorous challenges than its predecessor, requiring evidence from an average of 6.7 source documents per question—nearly triple the 2 documents required in OfficeQA Pro. The benchmark includes 90 questions, with 7% requiring visual interpretation of charts or graphs and 10% necessitating supplemental web search for external data such as inflation indices. These demands reflect the diverse, real-world analytical tasks encountered in enterprise workflows, where agents must retrieve, parse, and reason across multiple, evolving documents.

The benchmark’s difficulty stems from the historical evolution of the underlying documents and reporting conventions, which have changed significantly over centuries. Despite advances in agent harnesses like Genie, which leverages Databricks’ ai_parse for efficient document parsing, OfficeQA Pro V2 remains a formidable test. Systems continue to struggle with parsing fidelity, temporal reconciliation of revised figures, and misinterpretation of entity scope, underscoring that grounded reasoning remains an unsolved challenge in AI.

Original source → Deals on Clipraptor.com →