PatchBench

Contribute

Add your agent to the leaderboard.

Built a patching agent, or tried a new model with an existing one? Run it on the PatchBench tasks and send us the patches and reports from that run. The PatchBench repository has a harness and step-by-step instructions to get you started. After we check them, your agent appears in the table next to the 11 systems from the paper.

§ 01Rules

Tell us how you ran it.

Evaluate your agent on all 213 tasks. Keep a record of the setup you used, such as the exact model version, the reasoning settings, any time or budget limits, and any changes to the agent or its prompt.

The agent must not have web access during the run, since it could look up the historical fixes.

When you submit, also tell us how the agent and model should be named in the table, the version or commit you ran, a link to the project or paper, whether the code is open source, and the average cost per task.

§ 02Package

Bundle the output of your run.

Send us three things from your run. If you used the PatchBench harness, they are already in its out folder.

out/diffs/
The patch your agent produced for each task.
out/results/
The per-task verdicts and summary.csv written by analyze.
out/logs/infer/
The full trajectory of the agent on each task.

§ 03Submit

Send it to us as a GitHub issue.

Upload the archive anywhere we can download it from, then open an issue in the PatchBench repository with the link and the details of your entry. We rerun the validation stages on your patches before adding the row, and we will ask in the issue if anything is unclear. The same place works for questions about the benchmark, the harness or a result in the table.