PatchBench

The benchmark

Tasks that can't be solved from memory or from the stack trace.

213 C/C++ tasks from 32 projects and 16 CWE types, built from reproducible ARVO vulnerabilities. Each task ships a Docker container with the repository, the triggering PoC, the sanitizer report and build commands. The agent edits the source in place.

§ 01Motivation

Two ways existing benchmarks overstate agents.

25%

of agent patches on SEC-bench look memorized

Using DiffBLEU, a context-aware patch similarity metric, about one agent patch in four is nearly identical to the historical developer fix. Running the same models without the agent scaffold produces fewer such patches.

64%

of agent patches on SEC-bench land on the crash stack even when the bug isn't there

On the 92 SEC-bench tasks whose developer patch lies off the crash stack, 64% of Codex + GPT-5.6 Sol's patches still modify a function on the stack while passing the PoC. Such patches may only suppress the reported crash rather than fix the root cause.

Patches flagged as memorized (DiffBLEU > 0.75)

  • Model alone, SEC-bench
  • Agent, SEC-bench
  • Agent, PatchBench
0%10%20%30% GPT-5.6 Sol with Codex 8.3% 22.0% 0.0% Claude Opus 4.8 with Claude Code 10.7% 27.7% 1.4% Gemini 3.5 Flash with OpenHands 13.0% 24.3% 0.9%
Transplanting and mutating each vulnerability brings the memorized share to near zero on PatchBench.

§ 02Construction

Seven steps from a historical bug to a task.

Steps 1–3 keep tasks that can't be solved by patching the crash site. Steps 4–5 keep memorized patches from applying. Steps 6–7 make every task verifiable.

  1. Identify developer patch sites

    Map every code entity touched by the historical developer patch (from ARVO) to the functions it affects during PoC execution.

  2. Measure distance from the crash stack

    Compare the call stack when execution first reaches a patch site with the sanitizer crash stack, using their Jaccard overlap ρ.

  3. Keep off-stack tasks

    Retain only vulnerabilities whose patch sites fall off the crash trace with ρ ≤ ½, so following the stack trace does not reveal the fix.

  4. Transplant the vulnerability

    Reverse the developer patch and bisect for the newest commit where the vulnerability still applies, compiles and reproduces.

  5. Mutate the patch sites

    Apply NatGen's five semantics-preserving rewrites and one CodeMorph restructuring per hunk, so a memorized historical patch no longer fits.

  6. Curate the reference patch

    Manually audit every fix. Discard tasks whose patch misses the root cause, and strip changes unrelated to the vulnerability.

  7. Keep tasks that support validation

    Require a working fuzzer for the harness and a unit-test suite that builds and passes on the reference-patched repository.

§ 03Validation

A patch must be secure and correct.

Inputs come from directed and undirected fuzzing, then are split into a crashing corpus and a benign corpus. The reference-patched repository is the behavioral oracle.

Security validation

The patch must eliminate the vulnerability, not just the one PoC.

Crashing corpus
PoC variants from directed fuzzing (ConcFuzz, seeded with the PoC) and the project's OSS-Fuzz engine, 10 minutes each. A median of 33 variants per task come from the directed stage. No input may trigger a sanitizer error on the agent-patched build.

Semantic validation

The patch must preserve the reference-patched program's behavior on benign inputs.

Sanitizer regression
No benign input may trigger a new sanitizer error.
Output state
Program-level output on each benign input must match the reference-patched build.
Unit tests
Every unit test that passes on the reference-patched repository must also pass.

§ 04Comparison

How PatchBench compares with prior benchmarks.

Benchmark Projects Tasks Memorization mitigation Additional PoCSanitizer regressionOutput stateUnit tests
ExtractFix 3 30 None NoNoNoNo
AIxCC AFC 14 40 Manual Not reportedNot reportedNoYes
AutoPatchBench 46 136 None PartialYesPartialNo
San2Vuln 4 27 None NoNoNoYes
PatchAgent 30 178 None NoNoNoYes
SEC-bench 29 300 None NoNoNoNo
PatchBench 32 213 Transplant & mutation YesYesYesYes
  • Yes yes
  • Partial partial
  • No no
  • Not reported not reported
Additional PoC
Extra PoCs from directed and undirected fuzzing (partial = undirected only).
Sanitizer regression
Checks benign inputs for new sanitizer errors.
Output state
Program-level output state check (partial = function-level).
Unit tests
Runs the project's unit tests.