Study
When Execution Feedback Reveals the Expected Output: A Forensic Study of Answer Leakage in Agentic Code Repair
Zenodo, author preprint, not peer reviewed. DOI 10.5281/zenodo.21286873 · 10 July 2026
Coding agents are often given test-runner feedback and then credited with improving a repair. The preprint audits three experimental designs from the author's own research programme and finds that the feedback frequently exposed the expected output itself: in the primary QuixBugs arm, 156 of 169 consumed feedback messages (92.3 per cent) rendered the fixed expected-output field. The conclusion is that feedback which exposes a visible test assertion is causally confounded with verification, and that any claim of a verification benefit needs pass-or-fail-only feedback, a same-candidate-pool selection control, task-clustered inference and immutable transcript-to-row hashes. It is how NASCIA thinks about evaluation: an evaluation must test what its feedback reveals.
Read on Zenodo