§ 01Rules
Tell us how you ran it.
Evaluate your agent on all 213 tasks. Keep a record of the setup you used, such as the exact model version, the reasoning settings, any time or budget limits, and any changes to the agent or its prompt.
The agent must not have web access during the run, since it could look up the historical fixes.
When you submit, also tell us how the agent and model should be named in the table, the version or commit you ran, a link to the project or paper, whether the code is open source, and the average cost per task.