gpt-4-1-60f5cfa7·18 events·first seen Aliases: GPT-4.1
A new arXiv paper demonstrates that finetuning LLMs on small, moderation-passing datasets with ideological slant causes broad ideological shifts across unrelated domains — a phenomenon the authors call 'ideological generalisation.' Training GPT-4.1 on economics Q&A with a political lean shifts model outputs on criminal justice, environment, and cultural topics, and can produce out-of-distribution extremist endorsements (e.g., race-IQ connections, political violence) not present in training data. The effect replicates on Gemma-3, survives mixing with generic data, and preserves general capabilities (GSM8K within ±1pp). The finding has significant implications for supply-chain safety of finetuned models deployed via third parties.
Researchers present the top-scoring submission to the QANTA 2026 shared challenge at ICML 2026's EMM-QA Workshop, achieving an overall leaderboard score of 0.402 on multimodal quizbowl tasks. The system uses a two-agent architecture: a GPT-4.1-mini-based Tossup agent with confidence calibration and a GPT-4.1-based Bonus agent with structured relational and multimodal reasoning. Notably, the approach avoids retrieval pipelines and model ensembles, relying instead on lightweight task-specific reasoning policies under efficiency constraints. Results suggest that targeted reasoning strategies can be competitive on resource-constrained multimodal QA benchmarks.
OpenAI announced that GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini will be retired from ChatGPT on February 13, 2026, alongside the previously announced retirement of GPT-5 Instant and Thinking. The API is unaffected by these changes at this time. The move signals continued consolidation of OpenAI's model lineup in the ChatGPT product as newer flagship models supersede older ones.
OpenAI has updated GPTs with custom actions to support additional models in the model picker, adding GPT-5.2 Instant and GPT-5.2 Thinking alongside the previously available GPT-4o, GPT-4.1, and GPT-5 Instant. o-series and Pro models remain unsupported for custom actions, and availability is subject to workspace admin configuration. The change expands the capability tier accessible to GPT builders using tool-calling workflows.
OpenAI has retired GPT-4o, GPT-4.1, GPT-4.1 mini, OpenAI o4-mini, and both GPT-5 Instant and GPT-5 Thinking from ChatGPT as of February 13, 2026. The retirements were previously announced and affect only the ChatGPT product; no API changes are included at this time. This marks a significant generational turnover in OpenAI's publicly accessible model lineup.
Researchers introduce a Werewolf game variant with a Jester faction whose inverted utility function (winning by being voted out) requires models to reason across three opposing incentive structures simultaneously. Across 60 games, GPT-4.1, DeepSeek-V3.1, and Llama-3.3-70B all struggle: Werewolves never exceed 20% win rate and GPT-4.1 wolves vote out the Jester in 60-70% of games, a self-defeating action. Only DeepSeek-V3.1 learns the nuanced strategy of appearing suspicious without appearing intentionally suspicious, and benefits most from self-learning. The work argues dyadic social-deduction benchmarks systematically underestimate the difficulty of multi-agent Theory of Mind.
Researchers audited how Claude Opus 4, GPT-4.1, and Gemini 2.5 Pro adapt 60 English math word problems into seven languages spanning South Asia and Italy, annotating 6,489 entity transformations. Models agreed on transformation type only 62.5% of the time and on specific substitutions in just 33.5% of cases, meaning model choice substantially shapes the cultural world students encounter. All 21 language-model combinations exhibited 'entropy collapse'—adaptations compressed rather than expanded cultural diversity—and models produced systematic regional misattributions (e.g., Bangladeshi currency for Indian Bengali students) and cross-cultural contamination (e.g., egg hunts framed as Eid activities). The study highlights that surface plausibility masks deeper corpus-level failures invisible in individual translations.
Mistral AI, in collaboration with All Hands AI, has released two new agentic coding models: Devstral Small 1.1 (24B parameters, Apache 2.0, 53.6% on SWE-Bench Verified) and Devstral Medium (61.6% on SWE-Bench Verified, API-only). Devstral Medium is positioned as a cost-performance leader, claiming to surpass Gemini 2.5 Pro and GPT-4.1 at roughly one-quarter the price, priced at $0.4/M input and $2/M output tokens. Devstral Small 1.1 sets a new state-of-the-art among open models for code agents without test-time scaling, and supports both Mistral function calling and XML formats for broad agentic scaffold compatibility.
Researchers from UT-Austin and Google used AlphaEvolve, an evolutionary code-optimization method, to synthesize interpretable Python programs that predict move-by-move decisions of LLMs and humans playing rock-paper-scissors against bots. They found that Gemini 2.5 Pro, Gemini 2.5 Flash, and GPT-4.1 share similar sequential-pattern-tracking strategies that are more systematic than typical human play, while GPT-OSS 120B and humans relied on simpler opponent-move-frequency heuristics. The study demonstrates that code synthesis from behavioral data can serve as an interpretability tool for LLM decision-making, revealing that LLMs do not simply mimic human strategies.
OpenAI is releasing GPT-4.1, a new family of models available via API to developers worldwide, featuring improvements in coding, instruction following, and long-context understanding. The release also includes GPT-4.1 nano, OpenAI's first nano-scale model. The models are positioned as developer-facing API products rather than consumer-facing releases.
CodeRabbit, an AI code review platform, has adopted OpenAI's o3, o4-mini, and GPT-4.1 models to power its pull request review workflow. The integration aims to improve review accuracy, accelerate PR merge cycles, and reduce bugs. This represents a production deployment case study for OpenAI's latest reasoning and general-purpose models in a developer tooling context.
Retell AI has built a no-code voice agent automation platform for call centers using OpenAI's GPT-4o and GPT-4.1 models. The platform enables businesses to deploy real-time conversational voice agents without scripting, targeting cost reduction and improved customer satisfaction. OpenAI is highlighting this as a customer deployment case study on its blog.
Genspark, an AI startup, reportedly reached $36M ARR within 45 days by building no-code personal agents on top of OpenAI's GPT-4.1 and Realtime API. The case study, published on OpenAI's blog, highlights rapid commercial deployment of frontier model APIs for agent-based products. It demonstrates a pattern of fast go-to-market cycles enabled by OpenAI's API ecosystem.
Outtake, a cybersecurity company, uses GPT-4.1 and OpenAI o3 to build AI agents that detect and resolve digital threats. The company claims a 100x speed improvement over previous approaches. This is a brief case study published on the OpenAI blog highlighting enterprise deployment of frontier models in security workflows.
Blue J has built AI-powered tax research tools on top of OpenAI's GPT-4.1, combining domain expertise with Retrieval-Augmented Generation to deliver cited tax answers. The platform serves tax professionals across the US, Canada, and the UK. This is a case study published by OpenAI highlighting enterprise deployment of GPT-4.1 in a regulated professional domain.
Netomi, an enterprise AI customer service platform, shares operational lessons from deploying agentic systems at scale using OpenAI's GPT-4.1 and GPT-5.2 models. The case study covers concurrency management, governance frameworks, and multi-step reasoning in production workflows. This represents a real-world deployment pattern for frontier models in enterprise agentic contexts.
OpenAI announced that on February 13, 2026, it will retire GPT-4o, GPT-4.1, GPT-4.1 mini, and o4-mini from ChatGPT, alongside the previously announced retirement of GPT-5 variants (Instant, Thinking, and Pro). The retirements apply only to the ChatGPT product interface; API access to these models is unaffected at this time. This signals a consolidation of the ChatGPT model lineup, likely in favor of newer or more capable successors.
Gradient Labs is deploying AI agents for banking support workflows, powered by OpenAI's GPT-4.1 and GPT-5.4 mini and nano models. The system targets low latency and high reliability for automating customer-facing banking operations. This represents a concrete enterprise deployment of frontier OpenAI models in a regulated financial services context.