evidence-attribution-in-visual-document-understanding-without-coordinates-or-region-labels-7cc510aa·1 events·first seen Aliases: Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
Researchers present a study showing that replacing coordinate-based evidence attribution with a language interface—where vision-language models quote evidence verbatim and a multimodal retriever locates the quote—substantially reduces attribution hallucination in visual document understanding. Across six open VLMs on a bilingual CiteVQA subset, evidence recall rises from at most 8 points to 26–47 points and hallucination rate roughly halves. The authors extend this into a GRPO-based training recipe that uses retrieved region crops as reward signal, lifting an 8B model's strict attributed accuracy from 22.4 to 33.8 without any region-level labels. The work offers a practical path to grounded document QA without expensive spatial annotation.