understanding-reasoning-from-pretraining-to-post-training-6a14215b·1 events·first seen Aliases: Understanding Reasoning from Pretraining to Post-Training
A new arXiv preprint uses chess as a controlled testbed to study the relationship between pretraining and RL post-training across the full LLM training pipeline, spanning models from 5M to 1B parameters. Key findings: post-RL performance is well-predicted by pretraining loss, RL reward curve slope improves approximately linearly with pretraining tokens, and RL behaves differently on easy versus hard problems (amplifying already-preferred moves vs. surfacing near-absent correct moves). The predictive pattern transfers to a math-domain 1B model, suggesting the findings generalize beyond the chess testbed.