mm-issueloc-vl-emb-ebfc39bf·1 events·first seen Aliases: MM-IssueLoc-VL-Emb
Researchers introduce MM-IssueLoc, a benchmark of 652 issue-PR instances across 23 programming languages designed to evaluate whether AI systems actually use visual evidence (screenshots, error dialogs, UI states) when localizing bugs in code repositories. The benchmark provides file-level and function-level gold labels, paired text-only and with-image evaluation modes, and VCE-based diagnostics that convert images to structured text. Evaluation of current LLM-based and retrieval-based systems shows poor performance (best agent reaches 38.96 file Acc@5), and high scores on text-dominant SWE benchmarks do not transfer to multimodal localization. The work isolates visual evidence as an explicit evaluation variable, exposing a gap in existing software engineering agent benchmarks.