off-context-grpo-8c339441·1 events·first seen Aliases: Off-Context GRPO
Researchers introduce Off-Context GRPO (OC-GRPO), a variant of the GRPO reinforcement learning algorithm that addresses the 'zero learning signal' failure mode when models cannot solve hard problems. The method uses privileged guidance (e.g., solution prefixes) during training rollouts while applying an importance-corrected objective to avoid distributional mismatch with the unguided target. On standard mathematical reasoning benchmarks, OC-GRPO achieves a 3.9% absolute (13.8% relative) improvement over vanilla GRPO with negligible added cost.