swe-gym-c936cfc9·1 events·first seen Aliases: SWE-Gym
A new arXiv paper systematically audits SWE-bench Verified and finds that 13.6% of PR-Issue pairings exhibit misalignment across five patterns and eleven fine-grained scenarios, undermining the benchmark's validity as an LLM evaluation tool. The authors introduce PAIChecker, a three-phase multi-agent system for detecting such misalignment, achieving up to 92.12% binary accuracy on SWE-Gym and 91.67% on SWE-bench Multilingual. The finding is significant because SWE-bench is one of the most widely cited benchmarks for agentic coding capability, and systematic data quality issues could distort leaderboard rankings and capability claims.