identity-preference-optimization-0215e117·1 events·first seen Aliases: Identity Preference Optimization
Researchers propose a solver-guided LLM framework that selects among a library of MIP formulations for multi-warehouse inventory allocation on a per-instance basis, rather than applying a single fixed formulation. The system is trained via SFT, IPO preference optimization, and GRPO reinforcement learning using MIP solver evaluations as reward signals. Evaluated on real JD.com data, GRPO raises Hit Ratio@1 from 21.45% to 50.42% and achieves a 12.57 percentage point allocation accuracy gain over the incumbent baseline. The work demonstrates a practical pattern of using LLMs as meta-selectors over classical OR solvers in industrial logistics settings.