
gpt-4-1-mini-ce7dcba3·7 events·first seen Aliases: GPT-4.1 mini, GPT-4.1-mini
Researchers present the top-scoring submission to the QANTA 2026 shared challenge at ICML 2026's EMM-QA Workshop, achieving an overall leaderboard score of 0.402 on multimodal quizbowl tasks. The system uses a two-agent architecture: a GPT-4.1-mini-based Tossup agent with confidence calibration and a GPT-4.1-based Bonus agent with structured relational and multimodal reasoning. Notably, the approach avoids retrieval pipelines and model ensembles, relying instead on lightweight task-specific reasoning policies under efficiency constraints. Results suggest that targeted reasoning strategies can be competitive on resource-constrained multimodal QA benchmarks.
OpenAI announced that GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini will be retired from ChatGPT on February 13, 2026, alongside the previously announced retirement of GPT-5 Instant and Thinking. The API is unaffected by these changes at this time. The move signals continued consolidation of OpenAI's model lineup in the ChatGPT product as newer flagship models supersede older ones.
OpenAI has retired GPT-4o, GPT-4.1, GPT-4.1 mini, OpenAI o4-mini, and both GPT-5 Instant and GPT-5 Thinking from ChatGPT as of February 13, 2026. The retirements were previously announced and affect only the ChatGPT product; no API changes are included at this time. This marks a significant generational turnover in OpenAI's publicly accessible model lineup.
A new arXiv paper demonstrates that Handlebars templating's HTML auto-escaping—the default in Microsoft Semantic Kernel—provides uneven protection against structural role injection attacks, where attacker-controlled data carries chat role delimiters to forge higher-privilege turns. The authors conduct 5,760 trials across seven delimiter families, two attack objectives, and four models (GPT-3.5 Turbo, GPT-4o mini, GPT-4.1 mini, Claude Haiku 4.5), finding that HTML escaping neutralizes angle-bracket-based delimiters (ChatML, Llama-3, XML) but leaves colon- and Markdown-based families fully exposed. GPT-3.5 Turbo follows task-hijack instructions in 97% of raw and 91% of escaped trials; Claude Haiku 4.5 resists both objectives almost entirely. The paper concludes that HTML escaping cannot substitute for structural separation of instruction and data.
Mistral AI, in collaboration with All Hands AI, releases Devstral, an agentic LLM specialized for software engineering tasks under the Apache 2.0 license. The model achieves 46.8% on SWE-Bench Verified, surpassing prior open-source state-of-the-art by over 6 percentage points and outperforming larger models like DeepSeek-V3-0324 (671B) and Qwen3 232B-A22B under the same OpenHands scaffold. Devstral is small enough to run on a single RTX 4090 or a Mac with 32GB RAM, and is available via Mistral's API at $0.1/M input tokens, as well as on HuggingFace, Ollama, and other platforms. Mistral indicates a larger agentic coding model is in development.
Mistral AI has released Codestral 25.08, a code generation model update claiming +30% accepted completions, 50% fewer runaway generations, and improved FIM benchmark performance. The announcement also frames a full enterprise coding stack comprising Codestral (completion), Codestral Embed (code-specific retrieval), and Devstral (agentic workflows via OpenHands), all deployable on-prem or in VPC environments. Devstral Medium is reported to achieve 61.6% on SWE-Bench Verified, while Devstral Small (24B, Apache-2.0) reaches 53.6%. The pitch targets regulated industries blocked by SaaS-only competitors through self-hostable, air-gapped deployment options.
OpenAI announced that on February 13, 2026, it will retire GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini from ChatGPT, alongside the previously announced retirement of GPT-5 variants (Instant, Thinking, and Pro). The retirements apply only to the ChatGPT product interface; API access to these models is unaffected at this time. This signals a consolidation of the ChatGPT model lineup, likely in favor of newer or more capable successors.