os-shepherd-100k-54472cf7·1 events·first seen Aliases: OS-Shepherd-100K
Researchers introduce OSReward, a benchmark for evaluating vision-language model (VLM) judges on computer-using agent (CUA) trajectories, finding that even state-of-the-art models exhibit systematic leniency bias — mislabeling failed runs as successes. The study also releases OSReward-Hard and OSReward-Multi challenge sets, plus OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments. Built on this corpus, OS-Shepherd (9B and 35B) reward models are trained to match commercial judge quality at 30–60% lower cost. The work addresses a critical gap in scalable CUA evaluation infrastructure needed for both RL training and data curation.