hyperball-may-not-be-a-free-lunch-5f8a1aae·1 events·first seen Aliases: Hyperball May Not Be a Free Lunch
A new arXiv preprint investigates why Hyperball-style optimizers (which fix matrix parameter norms and normalize updates) outperform alternatives in large-scale training, finding that their advantage stems primarily from effective learning-rate schedule dynamics rather than an intrinsically superior update direction. The authors introduce an angular effective learning rate framework, decompose updates into radial and tangential components, and show that radial updates have limited direct effect on angular displacement. Experiments with MuonH and MuonWD reveal that careful learning-rate scheduling remains essential even under Hyperball constraints, contradicting the implicit assumption that norm-fixing eliminates scheduling sensitivity.