GitHub published ReviewBench on 5 October 2026, by Michelle Zhou and Alejandro Carderera de Diego. It is an offline benchmark for AI code-review agents, in research preview at review-bench.ai. The dataset, rubric, judge, and a self-serve runner are in the ReviewBench repository.
GitHub built the benchmark and uses it to evaluate Copilot code review, the product it sells. Related coverage of the 2 October API changelog is separate from this benchmark. The name is also taken: on 31 July 2026 LangChain published a different ReviewBench, 59 Harbor tasks from comments in its LangSmith repo. The figures below are GitHub’s set.
What the 219 pull requests are
GitHub says it analyzed 103.9 million GitHub pull requests to describe real review work. ReviewBench itself is 219 public pull requests from 187 public open-source repositories, across 19 languages. The post says language and repository-size distributions closely match GitHub overall. Pull-request size does not. GitHub says it deliberately weighted size toward the reviewable middle and tail, so tiny single-file changes are less dominant than they are on GitHub. The opening summary is looser. It says the set follows language, repository size, and size distribution. The method section is the one that states the size adjustment.
The corpus manifest opened with the repo contains 219 entries and 187 repository URLs. Its language field has 19 named languages plus one pull request labeled unknown. The largest counts in that file are TypeScript 68, Python 41, C# 25, Go 19, and JavaScript 15. Those counts are a reading of the manifest, not GitHub’s claim that the shape matches the 103.9 million.
How a finding becomes a label
GitHub says candidates come from human reviewers, issues inferred from later commits, deterministic analysis tools, and frontier models from several families. Overlaps are merged. A finding counts only if it is true, relevant, and non-trivial. The grader named in the post is Claude Sonnet 5, and GitHub says the rubric and judge are published. The benchmark site links that write-up to the repo’s methodology doc. Senior engineers who had not built the set then re-labeled every ground-truth finding, GitHub says, and agreed with it 96.6 percent of the time.
Grounded precision, recall, and F1 use only the golden labels. Augmented precision, recall, and F1 also let the judge accept a finding the golden set missed. GitHub says grounded recall is the cross-system headline, because augmented recall grows the denominator with whatever each agent adds. The post also describes an Fβ score that can favor recall or precision, plus slices by severity and category. The post’s severities are Critical, Medium, and Low. Its named categories are correctness, security, reliability, maintainability, and testing. The leaderboard feed opened the same day labels the top severity High, not Critical, and adds an API-design category.
Scores on GitHub’s own board
The post does not print a multi-system table. It says that offline ReviewBench changes, checked before A/B tests, “consistently pointed in the same direction” as later production. That is GitHub’s claim about Copilot code review. In the example, a lite-tier ensemble of several model runs was predicted to raise precision, recall, and comment volume, and to lower cost per review. Online, against the production control, addressed rate rose 8.0 percent, recall 13.6 percent, comment volume 61 percent, and cost per review fell 8.0 percent. GitHub defines addressed rate as the share of review comments that a model judges to have prompted a code change. It says ReviewBench predicted a 227 percent rise in critical comments, against 262 percent online.
The public leaderboard at review-bench.ai, fetched on 5 October 2026, had 28 rows. Rank is the site’s order. Rounded from the feed to one decimal, the first three featured rows were:
| Rank | Reviewer | Snapshot | Grounded precision | Grounded recall |
|---|---|---|---|---|
| 1 | Copilot Code Review, Balanced | 2026-10-01 | 87.8% | 26.0% |
| 2 | Devin AI | 2026-09-28 | 84.0% | 23.8% |
| 3 | Qodo | 2026-09-28 | 85.3% | 22.1% |
Each of those three is marked as three rounds. Copilot’s augmented F1 on that row is 49.7. Many remaining rows are Codex effort settings. GitHub’s own product is rank 1 on the board it publishes.
How an outside system is submitted
The post’s steps are: sign in with GitHub on the ReviewBench site; register a container image, a configuration, and your own model key, with GitHub supplying the judge; tune on a 25-pull-request test set; run the full 219 three times; then publish. Scores stay private until a maintainer approves them, and they are published only if they beat that agent’s current leaderboard score, or if it is the agent’s first entry.
Practical takeaway
The set is public, the manifest matches the 219 and 187 figures, and the scoring rules are specific enough to read before you trust a single F1. The comparison GitHub says to use across systems is grounded recall, not the augmented number. The production story is one vendor’s report that its offline runs moved with its online runs. The leaderboard is that same vendor’s ranking, with Copilot code review in first place.

The Campfire
No commentsNobody has pulled up a log by this one yet. Be the first to say what you make of it.
Held for the desk. It appears after a look.