25%
of agent patches on SEC-bench look memorized
Using DiffBLEU, a context-aware patch similarity metric, about one agent patch in four is nearly identical to the historical developer fix. Running the same models without the agent scaffold produces fewer such patches.