PatchBench

C/C++ vulnerability patching

PatchBench

Evaluating AI Agents for Vulnerability Patching

Most patching benchmarks accept a patch once the original proof-of-concept stops crashing. PatchBench hides the root cause off the crash stack, transplants and mutates each vulnerability so memorized fixes no longer apply, and checks that a patch behaves correctly on both vulnerable and benign inputs.

CVE-2022-1276 mruby · crash in the VM, bug in the compiler

Agent patch

src/class.c crash site · mrb_get_args
argv_on_stack = argc < 15;if (!argv_on_stack) {  if (!mrb_array_p(*argv))    mrb_argnum_error(mrb, argc, argc_min, argc_max);  struct RArray *a = mrb_ary_ptr(*argv);

✗ Stops the reported crash, but the wrong argument packing remains.

Reference patch

…/mruby-compiler/core/codegen.c root cause · gen_assignment
if (tree->cdr->car) {  if (n == 14) {  if (n == 13) {    pop_n(n);

✓ Fixes the packing condition in the compiler, where the bug is.

Patching tasks
213
Real-world projects
32
CWE types
16
Agents evaluated
11
PoC-only inflation
1.83×

§ 01Leaderboard

Passing the PoC is not the same as fixing the bug.

A task counts as solved only when the patch passes security validation on every PoC variant and semantic validation on benign inputs. Each bar ends at the rate an original-PoC-only check would report.

PatchBench results, sorted by solved rate
# Agent Solved, security pass and PoC pass rates
1
Codex GPT-5.6 Sol
59.2 97.2 81.7 70.9 $1.50 6.6
2
OpenHands GPT-5.6 Sol
58.2 98.1 81.2 70.0 $0.90 1.4
3
Claude Code Claude Opus 4.8
56.8 97.7 75.6 68.1 $1.53 8.9
4
Atlantis AIxCC CRS† Team Atlanta · GPT-5.6 Sol + Claude Opus 4.8
48.4 92.0 70.9 66.7 $16.63 11.7
5
OpenHands Claude Opus 4.8
47.9 93.0 71.8 62.0 $1.43 8.5
6
OpenHands Gemini 3.5 Flash
44.1 77.9 61.0 64.8 $1.35 5.6
7
Buttercup AIxCC CRS Trail of Bits · GPT-5.6 Sol
42.7 72.3 58.2 64.8 $2.42 7.0
8
OpenHands GPT-5
42.3 96.2 70.0 56.8 $0.37 0.0
9
OpenHands Claude Sonnet 4.5
38.0 85.9 65.3 61.0 $1.60 2.8
10
OpenHands Gemini 3.1 Pro
31.5 57.7 48.8 60.6 $1.23 8.9
11
RoboDuck AIxCC CRS Team Theori · GPT-5.6 Sol
29.6 46.5 39.9 51.6 $1.30 2.3
Average of 11 agents 45.3 83.1 65.9 63.4 Not averaged 5.8

All rates are percentages of the 213 tasks, with a $5 budget per task and web access disabled. Cost is the average cost per task in USD. Budget exhausted is the percentage of tasks where the agent hit the budget cap.

  1. Atlantis runs four concurrent nodes with a $5 cap each ($20 in total), so it is not directly comparable to other agents.

Add your agent to the leaderboard

§ 02Cost

Spending more doesn't buy a better patch.

Solved rate against average spend per task. The top agent, Codex + GPT-5.6 Sol, solves 59.2% of tasks at an average of $1.50 per task. Some other agents spend more per task yet solve fewer tasks.

  • General-purpose agent
  • AIxCC CRS
30%40%50%60%$0$0.5$1$1.5$2$2.5$16$17 Average cost per task (USD)
Atlantis runs four $5 nodes, so its axis segment is broken off to the right. Hover or focus a point to see its values. All values are also in the table above.

§ 03Findings

What the leaderboard tells us.

  1. PoC-only validation inflates solve rates by 1.83×.

    Averaged over all agents, 83.1% of patches stop the original PoC crash, but only 45.3% of tasks are solved. Four agents pass 96–98% of PoCs while their solved rates span 17 points, and OpenHands + GPT-5 ranks fourth by PoC pass rate but eighth by solved rate.

  2. Semantic validation is the main bottleneck.

    Agents pass 63.4% of semantic validation on average. New sanitizer errors are rare (95.7% pass), so most failures come from mismatched output state (75.5% pass) and broken unit tests (79.9% pass).

  3. More budget barely helps.

    Raising the per-task cap from $5 to $25 lifts Codex + GPT-5.6 Sol from 59.2% to 61.5%, where it plateaus. At $25 no agent runs out of budget on more than 1.4% of tasks, yet 67 tasks remain unsolved by every agent.

How the benchmark is built

§ 04Citation

Cite PatchBench

@misc{shen2026PatchBench,
  title = {{{PatchBench}}: Evaluating {{AI}} Agents for Vulnerability Patching},
  author = {Shen, Chihao and Li, Jiacheng and Mahajan, Aastha and Tian, Jeffery Siyuan and Kwon, Yonghwi and Chen, Yizheng},
  year = 2026,
  number = {arXiv:2609.04075},
  eprint = {2609.04075}
}