canon-1f6f2a9e·1 events·first seen Aliases: CANON
Researchers introduce CANON (Consensus-ANchored self-distillatiON), a label-free training method that converts majority-vote consensus from multiple sampled solutions into dense token-level supervision for LLM self-improvement. Unlike prior approaches that use consensus only as a filter, preference signal, or scalar reward, CANON conditions a frozen model snapshot on a consensus-reaching solution to supervise the model's own rollouts at every token. On mathematical and scientific reasoning benchmarks, CANON improves pass@1 by up to 12 points, outperforming label-free RL by 6 points at one-seventh the compute, and approaches performance of methods using gold labels. The method also transfers to held-out benchmarks when trained on pooled unlabeled data.