mle-bench-lite-a52a5c42·2 events·first seen Aliases: MLE Bench Lite, MLE-Bench Lite
Researchers introduce OpenMLE, an open full-stack system for recursive self-improvement (RSI) research in machine learning engineering, and post-train Frontis-MA1 (35B) as a meta-evolution agent on top of it. The system uses four atomic program-evolution operators (Draft, Improve, Debug, Crossover) trained via execution-grounded SFT and RL, then composed into long-horizon search. On MLE-Bench Lite under a constrained single-GPU budget, Frontis-MA1 improves Medal Average from 39.39% to 71.21% with OpenMLE-Evo-Max, reportedly exceeding GPT-5.5 + Codex and approaching GPT-5.6 Sol and the 2.8T Kimi K3. Model weights and the full OpenMLE stack are released publicly, targeting reproducible RSI research.
MiniMax released M2.7, a proprietary reasoning model that achieved 66.6% on MLE Bench Lite (tying Gemini 3.1) and 56.22% on SWE-Pro, priced at $0.30/$1.20 per million tokens, with the shift to proprietary marking a potential strategic pivot among Chinese AI labs away from open weights. Cursor released Composer 2, an agentic coding model built on a fine-tuned Kimi 2.5 (via Moonshot partnership), priced 86% cheaper than its predecessor and scoring 73.7 on SWE-bench Multilingual. Anthropic released Claude Code Channels, routing Telegram and Discord messages into local Claude Code sessions via MCP plugins, and separately filed a court response denying it has any backdoor or kill switch into military deployments of Claude. Microsoft announced MAI-Image-2, a text-to-image model ranking third on Arena.ai among research labs.