iso-an-rlvr-native-optimization-stack-f8ca8e42·1 events·first seen Aliases: ISO: An RLVR-Native Optimization Stack
Researchers introduce Isospectral Optimization (ISO), a framework that exploits 'spectral inheritance' in RLVR-trained language models — the observation that reward-driven adaptation changes singular frames while preserving base model weight spectra. ISO has two instantiations: ISO-Merger, a data-free method for combining specialist models without gradient updates or on-policy distillation, and ISO-Optimizer, which applies standard optimizers (AdamW, Muon) only to frame variables, achieving equivalent accuracy in roughly 2.7x fewer training steps on Qwen3-8B-Base. The work proposes a principled answer to the underexplored optimization layer between reward signals and weight updates in RLVR pipelines.