C/C++ vulnerability patching
PatchBench
Evaluating AI Agents for Vulnerability Patching
Most patching benchmarks accept a patch once the original proof-of-concept stops crashing. PatchBench hides the root cause off the crash stack, transplants and mutates each vulnerability so memorized fixes no longer apply, and checks that a patch behaves correctly on both vulnerable and benign inputs.
Agent patch
mrb_get_args
argv_on_stack = argc < 15;if (!argv_on_stack) { if (!mrb_array_p(*argv)) mrb_argnum_error(mrb, argc, argc_min, argc_max); struct RArray *a = mrb_ary_ptr(*argv);
✗ Stops the reported crash, but the wrong argument packing remains.
Reference patch
gen_assignment
if (tree->cdr->car) { if (n == 14) { if (n == 13) { pop_n(n);
✓ Fixes the packing condition in the compiler, where the bug is.
- Patching tasks
- 213
- Real-world projects
- 32
- CWE types
- 16
- Agents evaluated
- 11
- PoC-only inflation
- 1.83×
§ 01Leaderboard
Passing the PoC is not the same as fixing the bug.
A task counts as solved only when the patch passes security validation on every PoC variant and semantic validation on benign inputs. Each bar ends at the rate an original-PoC-only check would report.
| # | Agent | Solved, security pass and PoC pass rates | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 |
Codex
|
59.2 |
|
97.2 | 81.7 | 70.9 | $1.50 | 6.6 |
| 2 |
OpenHands
|
58.2 |
|
98.1 | 81.2 | 70.0 | $0.90 | 1.4 |
| 3 |
Claude Code
|
56.8 |
|
97.7 | 75.6 | 68.1 | $1.53 | 8.9 |
| 4 | 48.4 |
|
92.0 | 70.9 | 66.7 | $16.63† | 11.7 | |
| 5 |
OpenHands
|
47.9 |
|
93.0 | 71.8 | 62.0 | $1.43 | 8.5 |
| 6 |
OpenHands
|
44.1 |
|
77.9 | 61.0 | 64.8 | $1.35 | 5.6 |
| 7 |
Buttercup AIxCC CRS
|
42.7 |
|
72.3 | 58.2 | 64.8 | $2.42 | 7.0 |
| 8 |
OpenHands
|
42.3 |
|
96.2 | 70.0 | 56.8 | $0.37 | 0.0 |
| 9 |
OpenHands
|
38.0 |
|
85.9 | 65.3 | 61.0 | $1.60 | 2.8 |
| 10 |
OpenHands
|
31.5 |
|
57.7 | 48.8 | 60.6 | $1.23 | 8.9 |
| 11 |
RoboDuck AIxCC CRS
|
29.6 |
|
46.5 | 39.9 | 51.6 | $1.30 | 2.3 |
| Average of 11 agents | 45.3 | 83.1 | 65.9 | 63.4 | Not averaged | 5.8 |
All rates are percentages of the 213 tasks, with a $5 budget per task and web access disabled. Cost is the average cost per task in USD. Budget exhausted is the percentage of tasks where the agent hit the budget cap.
- Atlantis runs four concurrent nodes with a $5 cap each ($20 in total), so it is not directly comparable to other agents.
§ 02Cost
Spending more doesn't buy a better patch.
Solved rate against average spend per task. The top agent, Codex + GPT-5.6 Sol, solves 59.2% of tasks at an average of $1.50 per task. Some other agents spend more per task yet solve fewer tasks.
- General-purpose agent
- AIxCC CRS
§ 03Findings
What the leaderboard tells us.
-
PoC-only validation inflates solve rates by 1.83×.
Averaged over all agents, 83.1% of patches stop the original PoC crash, but only 45.3% of tasks are solved. Four agents pass 96–98% of PoCs while their solved rates span 17 points, and OpenHands + GPT-5 ranks fourth by PoC pass rate but eighth by solved rate.
-
Semantic validation is the main bottleneck.
Agents pass 63.4% of semantic validation on average. New sanitizer errors are rare (95.7% pass), so most failures come from mismatched output state (75.5% pass) and broken unit tests (79.9% pass).
-
More budget barely helps.
Raising the per-task cap from $5 to $25 lifts Codex + GPT-5.6 Sol from 59.2% to 61.5%, where it plateaus. At $25 no agent runs out of budget on more than 1.4% of tasks, yet 67 tasks remain unsolved by every agent.
§ 04Citation
Cite PatchBench
@misc{shen2026PatchBench,
title = {{{PatchBench}}: Evaluating {{AI}} Agents for Vulnerability Patching},
author = {Shen, Chihao and Li, Jiacheng and Mahajan, Aastha and Tian, Jeffery Siyuan and Kwon, Yonghwi and Chen, Yizheng},
year = 2026,
number = {arXiv:2609.04075},
eprint = {2609.04075}
}