entrap-vl-d17b897c·1 events·first seen Aliases: ENTRAP-VL
Researchers introduce ENTRAP-VL, a manually curated 1,500-item benchmark designed to probe contextual entrainment in vision-language models (VLMs) — the tendency of models to let irrelevant or false auxiliary context distort their outputs. The benchmark is taxonomically structured across two axes (context association and truth relationship) and two modality streams (textual and visual entrainment), covering eight categories. The paper argues that entrainment in VLMs is a substantively new dual-modality phenomenon not captured by porting unimodal text benchmarks, and the dataset will be released publicly.