GitHub's blog post ReviewBench (Oct 5, 2026) launches an offline AI code-review benchmark as a research preview. It has 219 PRs from 187 open-source repos in 19 languages, sampled to match the distribution of 103.9M GitHub PRs. The golden set combines human reviewers, several frontier LLMs and static analysis. Claude Sonnet 5 is the judge, and senior engineers agreed with its labels 96.6% of the time. On the leaderboard, grounded recall is 26.0% for Copilot Code Review (Balanced), 23.8% for Devin, 22.1% for Qodo and 9.6% for Cursor. The dataset, judge prompt and runner are MIT-licensed on GitHub.

Key Takeaways

  • ✓Scale: 219 PRs from 187 open-source repos in 19 languages, sampled to match the distribution of 103.9M GitHub PRs (blog)
  • ✓Trust: senior engineers independently re-labeled every golden finding and agreed with the benchmark 96.6% of the time; Claude Sonnet 5 is the judge
  • ✓Leaderboard grounded recall: Copilot Balanced 26.0% (87.8% precision) > Devin 23.8% > Qodo 22.1% > Codex GPT-5.6 Sol Ultra 19.9% > Cursor 9.6%
  • ✓Offline matched online: in a multi-model ensemble experiment, the online A/B showed +8.0% precision, +13.6% recall and −8.0% cost per review; critical comments were predicted at +227% offline and measured at +262% online
  • ✓To submit: bring your own container image and model key, run the 25-PR test set, then the full 219 PRs over 3 rounds. The repo review-bench/ReviewBench is MIT-licensed
GitHub open-sources ReviewBench, a 219-PR benchmark for AI code review: Copilot, Devin and Qodo lead; Cursor's grounded recall is 9.6%
🖼️Official Media
Click to view high-res
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Context

AI code reviewers such as Copilot, Claude Code, Devin, Qodo and Greptile are now common on pull requests, but each vendor reports its own metrics. GitHub's ReviewBench (research preview, Oct 5, 2026) is meant to be a shared, reproducible yardstick.

How it works

The corpus has 219 PRs from 187 open-source repos in 19 languages, sampled to match the distribution of 103.9M GitHub PRs and weighted toward substantive multi-file changes. Candidate findings come from human reviewers, issues inferred from authors' follow-up fix commits, static analysis and several frontier LLMs. They are semantically deduplicated, then checked against one rubric by a Claude Sonnet 5 judge. Senior engineers independently re-labeled the golden set and agreed 96.6% of the time. Grounded precision, recall and F1 count only known labels. Augmented metrics also give credit for valid new findings. Results can be sliced by severity and category, and β in the Fβ score can be tuned to favor recall or precision.

Numbers

Grounded recall on the leaderboard: Copilot Code Review Balanced 26.0% (87.8% precision), Devin 23.8%, Qodo 22.1%, Codex GPT-5.6 Sol Ultra 19.9%, Cursor 9.6%. Precision sits between 84% and 90% for every reviewer, so the real gap is recall. GitHub built the benchmark and its own product ranks first, so treat the ranking with care until third parties reproduce it. GitHub also reports that offline results predicted the direction of production A/B tests: +13.6% recall and −8.0% cost per review for a multi-model ensemble.

Getting started

Sign in at review-bench.ai and register a container image (referenced by digest), your configuration and your own model key. Follow ONBOARDING.md and the examples/codex-cli reference agent. Run the 25-PR test set first, then the full 219 PRs over 3 rounds. A maintainer reviews each submission before it is published.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.