activevision-ad99b6d8·1 events·first seen Aliases: ActiveVision
Researchers introduce ActiveVision, a 17-task benchmark designed to measure whether multimodal LLMs can perform active visual observation — iteratively redirecting attention based on intermediate hypotheses rather than processing a single static image. Frontier models collapse catastrophically: GPT-5.5 at maximum reasoning effort scores only 10.6% and zeros out on 11 of 17 tasks, while Claude Fable 5 scores 3.5%, versus 96.1% average for human participants. The gap persists even when models write and execute their own vision code, because detecting code failures itself requires the active perception the models lack. The results suggest a fundamental architectural gap between current MLLMs and human closed-loop visual cognition.