groc-po-760caaaa·1 events·first seen Aliases: Groc-PO
Researchers introduce Groc-PO, a preference optimization framework for multimodal LLMs that applies stage-specific supervision across object grounding, contextual grounding, and grounded reasoning stages, rather than only at the final-answer level as in standard DPO. The method is paired with a new dataset (GCPD) organized around these three grounding stages. Experiments show improvements over standard DPO and other baselines on hallucination mitigation and faithful reasoning, addressing the credit-assignment problem where early grounding errors propagate through reasoning chains.