mirror-a91c545b·1 events·first seen Aliases: MIRROR
Researchers introduce MIRROR (Modality-Informed Reciprocal Reasoning Optimization), a reinforcement learning approach that improves vision-language model reasoning by exploiting inconsistencies across different problem representations. The method constructs ODA-Data, a paired multimodal geometry dataset with text-dominant, image-dominant, and combined views, then uses the best-performing view as a teacher to train weaker views via reverse-KL divergence. MIRROR outperforms standard RL baselines on geometry reasoning benchmarks and produces more consistent cross-modal behavior. The work addresses a known gap between LLM and VLM reasoning capabilities on problems that have equivalent multi-modal representations.