OFICIAL GitHub Blog

ReviewBench: An open benchmark for AI code review

What happened
Based on GitHub Blog · Oct 05, 2026

GitHub introduces ReviewBench, an open benchmark for evaluating AI code review agents using real pull requests and multi-source ground truth to improve reliability and production alignment.

Key points
·
ReviewBench uses 219 public pull requests from 187 repositories across 19 languages to model real GitHub review workloads.
·
Senior engineers validated ReviewBench’s ground truth with 96.6% agreement on re-labeled findings before release.
·
ReviewBench’s six metrics include grounded recall and augmented recall to capture newly discovered issues in reviews.
Key numbers
·
ReviewBench models its pull request language, repository size, and size distribution after over 100 million real GitHub pull requests.
·
6% agreement with ReviewBench’s judgments.

ReviewBench is a new open benchmark designed to evaluate AI code review agents by leveraging representative GitHub pull requests and multi-source ground truth. It aims to address inconsistencies in measuring reviewer quality, where some systems catch more issues while others generate noise or focus on critical problems. The benchmark provides a rigorous and reproducible evaluation methodology that reflects real-world code review diversity, including severity, category, and precision-recall tradeoffs. It is built to support teams in building and improving code review agents with an offline signal that predicts production performance.

ReviewBench models its pull request language, repository size, and size distribution after over 100 million real GitHub pull requests. It includes 219 public pull requests from 187 open source repositories across 19 languages, closely matching GitHub-wide distributions. The benchmark adjusts pull request size to reduce overrepresentation of trivial changes while preserving substantive, multi-file reviews where review quality matters most. Every finding is labeled for severity and category, enabling tailored analysis by user needs.

The benchmark introduces six metrics in two families to evaluate reviewer performance, including grounded recall and augmented metrics that account for newly discovered issues. Senior engineers independently validated the ground truth by re-labeling findings, achieving 96.6% agreement with ReviewBench’s judgments. The benchmark is versioned to ensure consistent comparisons and includes published validation methodology and known limitations for transparency.

ReviewBench is available as a research preview on its website, allowing users to explore the dataset, compare agents, and submit their own systems for evaluation. GitHub used ReviewBench to evaluate Copilot code review, improving its ability to predict production performance and prioritize changes. Offline evaluations with ReviewBench have consistently aligned with later A/B test results, demonstrating its reliability as an early signal for product improvements.

Original source → Deals on Clipraptor.com →