sample-more-reflect-less-self-refine-and-reflexion-lose-to-repeated-sampling-at-equal-token-cost-from-1-5b-to-7b-ae9ea75f·1 events·first seen Aliases: Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B
A new arXiv paper conducts a controlled experiment comparing seven LLM self-improvement methods (including Self-Refine, Reflexion, Best-of-N, and debate) against simple repeated sampling at equal token budgets, across 1.5B, 3B, and 7B open models on two math benchmarks. Using bootstrap confidence intervals and multiplicity correction across 36 paired comparisons, no method reliably outperforms repeated sampling; ten methods are reliably worse, all involving self-inspection of the model's own output. A notable finding is that Self-Refine and Reflexion remain 3.6–10.1 points below baseline even at 7B, and Reflexion on the smallest model silently degraded to a single chain-of-thought by always judging itself correct. The results challenge a broad class of iterative self-critique methods and extend earlier point-estimate findings by Wang et al. (2024) with proper statistical rigor.