klear-reasoner-8b-sft-c8a13aa6·1 events·first seen Aliases: Klear-Reasoner-8B-SFT
A new arXiv preprint introduces Lightning OPD 2.0, a method for on-policy distillation (OPD) that addresses style bias when the SFT data generator and distillation teacher are different models. The approach uses rollout-level cross-fitting to estimate and subtract a 'style residual' from teacher-reference disagreement before constructing token-level updates. Starting from Klear-Reasoner-8B-SFT, the method achieves 82.4% on AIME 2024 and 63.0% on LiveCodeBench v5, outperforming the original Lightning OPD in cross-teacher settings. The work relaxes a key practical constraint in distillation pipelines by decoupling SFT data generation from the distillation teacher.