Researchers propose a solver-guided LLM framework that selects among a library of MIP formulations for multi-warehouse inventory allocation on a per-instance basis, rather than applying a single fixed formulation. The system is trained via SFT, IPO preference optimization, and GRPO reinforcement learning using MIP solver evaluations as reward signals. Evaluated on real JD.com data, GRPO raises Hit Ratio@1 from 21.45% to 50.42% and achieves a 12.57 percentage point allocation accuracy gain over the incumbent baseline. The work demonstrates a practical pattern of using LLMs as meta-selectors over classical OR solvers in industrial logistics settings.
This paper introduces an agentic framework where an LLM acts as an operations research expert, translating natural-language user prompts into structured updates ('patches') to deployed optimization models and selecting appropriate re-optimization techniques from a toolbox. The toolbox leverages primal information—historical solutions, valid inequalities, solver configurations, and metaheuristics—to accelerate re-optimization while preserving solution quality. Experiments on supply chain re-optimization and university exam scheduling demonstrate computational efficiency gains and improved interpretability through patch-based model modifications. The framework aims to reduce dependence on OR experts for maintaining dynamic decision-support systems.
Researchers introduce Agora, a framework that uses an incentive-compatible auction mechanism to dynamically route reasoning subtasks to the most capable expert models or tools, rather than relying on coarse-grained function matching. Agents bid based on 'rectified competence' to prevent overconfident solvers from capturing critical logic steps. Evaluations across five benchmarks show improvements over single-model, routing, and cascade baselines, with a controllable cost-quality trade-off via a single auction parameter.
Researchers from MERL propose LLawCo (Learning Laws of Cooperation), a framework that enables embodied LLM-based agents to autonomously align with partners and task objectives in decentralized, partially observable environments. Agents reflect on past failures to extract misaligned behavioral patterns and derive high-level behavioral laws (e.g., 'Talk when necessary', 'Wait for partner'), which are incorporated into reasoning via supervised fine-tuning. The authors also introduce PARTNR-Dialog, a new large-scale multi-agent communicative planning benchmark, and report average success rate improvements of 4.5% on PARTNR-Dialog and 6.8% on TDW-MAT over state-of-the-art open-source communicative agent frameworks across four backbone LLMs.
Researchers introduce a scalable benchmark for evaluating LLM agents on cooperative joint decision-making tasks where agents must exchange information under partial and asymmetric observations to reach a shared decision. A systematic evaluation of representative LLMs finds that state-of-the-art models still struggle with complex deliberative collaboration, failing in either information alignment or downstream reasoning even with external mathematical tools. Diagnostic analysis also reveals that deliberation can enable reflection and error correction, sometimes outperforming centralized baselines, offering a nuanced picture of multi-agent LLM capabilities.
OptiAgent is a multi-agent LLM framework that converts natural language descriptions of Operations Research problems into solver-ready mathematical formulations and executable code. The architecture uses dedicated agents for extracting decision variables and constraints, with a multi-loop validation system featuring four specialized feedback mechanisms targeting distinct failure modes. The system claims state-of-the-art performance on 3 of 4 benchmarks spanning LP, MILP, and Nonlinear Programming tasks, while also improving transparency through auditable agent reasoning.
This paper investigates multi-objective prompt optimization for LLM-as-judge systems, testing five decomposition modes of textual gradient optimizers across varying levels of cross-task information sharing. In 6 of 10 configurations, optimization fails to improve over the initial prompt, with gradient specificity dropping 59% when multiple criteria are processed jointly. The authors identify two separable failure modes: gradient dilution at optimization time and instruction interference at inference time. These findings constrain the design space for customizing LLM judges via textual feedback across multiple evaluation criteria simultaneously.
A new arXiv preprint investigates whether LLMs can replicate systematic human decision-making biases in route choice scenarios without explicitly specifying cumulative prospect theory (CPT) parameters. The authors design a behavioral evaluation framework and find that LLMs reproduce non-rational choice patterns consistent with CPT effects under uncertainty. The findings suggest LLMs could serve as a scalable substitute for survey-based parameter calibration in agent-based behavioral simulations, addressing a longstanding bottleneck in large-scale human decision modeling.
A preprint uses an open problem from EC 2025 as a testbed to evaluate AI-assisted research workflows in economics and computer science. The study examines whether human intuition in prompts, multi-turn interaction, and LLM capability compare favorably to a first-year PhD student's contributions. Key findings: human intuition in prompts improves LLM 'taste', multi-turn workflows help when encouraging ambitious steps, and the LLM performs slightly below the first-year PhD student on the same problem. The work contributes empirical evidence on the practical utility and limits of LLMs as research collaborators in formal theory domains.