October 6, 2026
Compare AI reviewers on shared tasks: ReviewBench brings together 219 pull requests
GitHub proposes measuring how many important bugs an AI reviewer finds and how many unnecessary comments it leaves.

GitHub
@github
How can you tell whether an AI reviewer finds important problems in code without adding unnecessary comments? ReviewBench is a new open benchmark: we built it based on an analysis of 103.9 million pull requests on GitHub and included 219 pull requests across 19 programming languages. Connect your code review agent, evaluate it, and submit the results ⬇️ https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/?utm_source=x-promoting-reviewbench-blog-article&utm_medium=social&utm_campaign=reviewbenchmark-oct-2026
· 7.9K views
ReviewBench tests AI review agents on real code changes. GitHub assembled 219 pull requests across 19 programming languages and described the benchmark on Oct 5, 2026.
A sample from real-world work. GitHub analyzed 103.9 million pull requests to select tasks spanning languages, repository sizes, and change sizes. Those millions informed the selection; the agent itself runs through a set of 219 tasks.
Bugs and noise. When evaluated against a fixed list of bugs, a reviewer earns credit for matching that list. ReviewBench also checks new findings from the agent to credit useful comments that were missing from the original list. The results table lets you compare critical bug detection and the proportion of correct comments separately.
GitHub already uses ReviewBench to evaluate reviews in GitHub Copilot. To test your own agent on the ReviewBench site, sign in with GitHub, register a container image and configuration, then add your model key. You start with a trial set of 25 pull requests, after which you can run the full evaluation.
Original source: [GitHub, Oct 5, 2026](https://github.blog/ai-and-ml/github-copilot/reviewbench-an-open-benchmark-for-ai-code-review/).
Your agent's results remain private until a ReviewBench maintainer reviews and approves the submission.
Source