looping-is-not-reliability-state-bound-evidence-and-typed-revision-contracts-for-agentic-code-repair-abf1ca5f·1 events·first seen Aliases: Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
A new arXiv preprint studies the reliability gap in generate-test-revise loops used by coding agents, finding that forced revision cycles cause current correctness to drop from 0.820 to 0.673 even as ever-correct rises to 0.847. Controlled experiments with 2,430 branches show stale traces harm 34/135 correct starts versus 4/135 with current traces, a statistically significant 22.2-point increase. The authors formalize the problem by separating admission, preservation, and certification concerns, then derive a typed loop contract with a mechanically enforceable reference implementation that binds verifier evidence to exact code states. The work is framed explicitly as a specification artifact rather than a claim of improved repair competence.