llama-3-2-1b-instruct-4631b55e·1 events·first seen Aliases: Llama-3.2-1B-Instruct
This arXiv paper proposes Output Reset (OR), a smooth one-sided saturation rule to replace the clipped surrogate objective in PPO and GRPO during LLM post-training. Experiments on Llama-3.2-1B-Instruct with the Anthropic hh-rlhf dataset show PPO-OR achieves a 0.305 higher mean reward-model score than PPO-clip under GAE, while GRPO-OR shows reduced variance but no reward gain at group size G=2. The work identifies a meaningful behavioral difference between the two optimization regimes but leaves open whether larger group sizes change GRPO-OR's effectiveness.