a-learning-rate-gated-failure-of-grpo-in-a-small-language-and-vision-language-model-web-agent-8120ee4b·1 events·first seen Aliases: A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent
A controlled ablation study across 18 runs tests whether GRPO reinforcement learning adds capability to 4B–8B scale language and vision-language model web agents on top of a strong supervised baseline. The result is a credible null: GRPO does not improve success rates when the supervised model has largely mastered the task distribution, and moderate-to-high learning rates actively degrade text-track performance. The authors identify the mechanism — GRPO only helps when sampling headroom exists (sampled policy succeeds more than greedy), and failure modes dissociate into attention/MLP degradation versus full collapse regimes. Effective rank in late layers tracks capability at 4B but not 8B, flagging a scale-dependent coupling.