post-training-shifts-confidence-a-three-stage-analysis-of-how-sft-rl-and-opd-shape-pre-intra-and-post-cot-calibration-85e72837·1 events·first seen Aliases: Post-Training Shifts Confidence: A Three-Stage Analysis of How SFT, RL, and OPD Shape Pre-, Intra-, and Post-CoT Calibration
A new arXiv paper introduces a three-stage calibration framework analyzing how supervised fine-tuning (SFT), reinforcement learning (RL), and on-policy distillation (OPD) shape model confidence before, during, and after chain-of-thought reasoning. The authors find that each post-training method produces distinct calibration profiles at different reasoning stages, and that RL confidence becomes informative only after a path-commitment phase while OPD confidence degrades later. They propose PosConf, a position-aware confidence strategy that selectively uses confidence from reliable relative-position intervals, improving RL answer aggregation by 6.1 points over majority voting and OPD early stopping by up to 4.3 points.