spatio-temporal-video-grounding-6cd837d7·1 events·first seen Aliases: Spatio-Temporal Video Grounding
Researchers introduce AnyGroundBench, a benchmark designed to evaluate Vision-Language Models on Spatio-Temporal Video Grounding (STVG) across five specialized domains: animal behavior, industry, sports, surgery, and public security. Unlike existing benchmarks that focus on zero-shot general-domain evaluation, AnyGroundBench includes dedicated training subsets to measure domain adaptability and in-context learning. Evaluation of 15 state-of-the-art VLMs reveals significant failures in both zero-shot and ICL-based adaptation, exposing weaknesses in spatio-temporal reasoning under domain shift.