pope-popperian-placebo-controlled-evaluation--57d250c4·1 events·first seen Aliases: PoPE (Popperian Placebo-controlled Evaluation)
A preregistered arXiv preprint introduces PoPE (Popperian Placebo-controlled Evaluation), a methodology for rigorously measuring whether execution error feedback actually improves code generation in frozen small LLMs (0.5–1.5B parameters). The study uses channel-specific placebos—ablating error content while preserving scaffold form—to test both prompt-based and adapter-based self-repair. Results across both channels failed to confirm content-attributable superiority over placebos or baselines, suggesting that apparent self-repair gains in prior work may reflect form rather than semantic error content. The paper argues that writing oracle-derived representations back into generation state replaces testing with conditioning.