test-time-training-for-modality-order-consistency-in-vision-language-models-9721b0f0·1 events·first seen Aliases: Test-Time Training for Modality Order Consistency in Vision-Language Models
Researchers find that vision-language models consistently perform better when images are presented before questions rather than after, a semantically irrelevant prompt-order artifact that persists across three models and three benchmarks. The authors propose an order-consistent test-time training method that closes this gap and also improves the stronger image-first baseline. Activation patching localizes the failure to a narrow mid-network region where representations diverge between orderings, framing modality-order sensitivity as a circuit-level failure. The work contributes both a diagnostic finding about VLM robustness and a lightweight mitigation technique.