evidence-backed-video-question-answering-eee2f28f·1 events·first seen Aliases: Evidence-Backed Video Question Answering
Researchers from Salesforce AI Research propose Evidence-Backed Video Question Answering (E-VQA), a task requiring Video LLMs to jointly produce semantic answers and precise spatio-temporal evidence via dense tracked object segmentation masklets. They introduce ST-Evidence, the first human-verified pixel-level grounding benchmark for video QA, and ST-Evidence-Instruct, a 160k-scale training dataset. Fine-tuning on this data yields substantial gains over size-matched baselines (+27.2 t-mean, +13.8 J&F on a 7B model), revealing a critical gap between QA accuracy and true visual perception that scaling alone does not close.