OFICIAL Databricks Newsroom Gadgets · Aug 06, 2026

Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning

In brief · 4 sentences
Based on Databricks Newsroom · Aug 06, 2026

Databricks released OfficeQA Pro V2, a benchmark testing AI agents' ability to generalize grounded reasoning across unfamiliar enterprise document collections, using 120,000 pages of U.S. Treasury accounting records.

Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning
Databricks Newsroom — Databricks
Key points
·
Main topic: introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning.
·
Category affected: gadgets and hardware.
·
Figures mentioned: 120,000, 11, 90.
·
The information comes from an official source.
·
The next step is to watch availability, pricing and real-world impact.

The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.

Databricks introduced OfficeQA Pro V2 to assess whether AI systems can generalize grounded reasoning beyond a single document collection, addressing a gap identified in the original OfficeQA benchmark released seven months prior. The new benchmark evaluates agents on 90 analytical questions using approximately 120,000 pages from the U.S. Treasury’s Accounts of Receipts and Expenditures, a dataset newly released for the U.S. Treasury’s 250th anniversary. The benchmark was first used in the Databricks Grounded Reasoning Cup, where 11 academic teams, supported by OpenAI, Anthropic, and Google DeepMind, developed agents evaluated on unseen data.

Out-of-the-box frontier agents achieved an average accuracy of 37.5% on OfficeQA Pro V2, while agents specifically developed for the competition averaged 41.1%, with the top team reaching 63.3%. Databricks’ Genie agent, using the same underlying models, improved accuracy by an average of 24.0 percentage points, with its strongest configuration reaching 60%. The results indicate that while grounded reasoning remains challenging, targeted agent harnesses can significantly enhance performance without requiring new model capabilities.

Baseline evaluations of five frontier models using their default harnesses showed an average accuracy of 26.0%, with performance varying widely and higher costs not consistently correlating with better results. For example, Sonnet 5 on Claude Code scored 15.6% at $5.01 per rollout, while GPT-5.6 Sol on Codex achieved 33.3% at $4.70. Genie’s use of Databricks’ ai_parse to pre-process documents reduced costs and improved accuracy, demonstrating that harness improvements can yield substantial gains in efficiency and performance.

OfficeQA Pro V2 introduces a more demanding test of grounded reasoning by requiring evidence from an average of 6.7 source documents per question, compared to 2 in the original OfficeQA Pro. The benchmark also includes 7% visual understanding questions and 10% requiring external web search, reflecting real-world enterprise challenges. Despite these advances, systems continue to struggle with parsing fidelity, temporal reconciliation, and entity scope misinterpretation, underscoring the need for further progress in enterprise AI capabilities.

Original source → Deals on Clipraptor.com →
Extracted signals · detected in the story
Introducing OfficeQA Pro V2New BenchmarkEnterprise Grounded-Reasoning. Meet OfficeQA ProU.STreasuryTodayOfficeQA Pro V2SevenOfficeQASince120,00011904.85