OpenAI has developed GPT-Red, an adversarial LLM designed to act as a sparring partner for its production models, stress-testing their defenses against cyberattacks. The system was used in training GPT-5.6, OpenAI's latest flagship model released last week, which the company claims is its most robust release to date. GPT-Red automates red-teaming at scale, representing a shift toward using AI systems to evaluate and improve AI safety properties.
OpenAI has published details on GPT-Red, an automated red teaming system that uses self-play to iteratively improve AI robustness against adversarial prompts, safety failures, and prompt injection attacks. The system is designed to enable self-improvement loops for alignment and safety testing without requiring constant human red teamers. This represents a notable step toward scalable automated safety evaluation infrastructure.
OpenAI is expanding its Trusted Access for Cyber program with two new models: GPT-5.5 and GPT-5.5-Cyber, a specialized variant aimed at cybersecurity applications. The program provides verified defenders with access to these models to accelerate vulnerability research and protect critical infrastructure. This represents a continuation of OpenAI's strategy of releasing domain-specialized model variants with controlled access tiers for sensitive use cases.
OpenAI is expanding its Trusted Access for Cyber program, introducing a specialized model called GPT-5.4-Cyber to vetted cybersecurity defenders. The program aims to provide advanced AI capabilities to legitimate security professionals while strengthening safeguards against misuse. This represents a structured approach to deploying frontier AI in sensitive security contexts with access controls.
This Hugging Face blog post introduces red-teaming as a safety evaluation methodology for large language models, explaining how adversarial testing can surface harmful outputs, biases, and failure modes before deployment. It covers techniques for systematically probing LLMs to elicit problematic behaviors and discusses the role of red-teaming in responsible AI development. The post serves as an educational overview aimed at practitioners working on LLM safety.
OpenAI is launching an open call for a Red Teaming Network, inviting domain experts to participate in ongoing safety evaluations of its models. The initiative aims to build a structured community of external red teamers who can help identify risks and failure modes across OpenAI's model releases. This represents a formalization of OpenAI's external adversarial testing program beyond one-off pre-release red teaming exercises.
OpenAI has published a system card update for GPT-5.2, the latest model family in the GPT-5 series. The safety mitigation approach is described as largely consistent with the prior GPT-5 and GPT-5.1 system cards. Training data sources follow the same pattern as other OpenAI models: publicly available internet data, third-party partnerships, and user/researcher-generated content.
OpenAI announced GPT-5.6, a new flagship model positioned as offering more intelligence per token, stronger performance per dollar, and increased capability for demanding tasks. The release is framed around efficiency and scalability alongside raw capability. This supersedes GPT-5.5 as OpenAI's current flagship.
OpenAI announced a preview of three vision-language models — GPT-5.6 Sol, Terra, and Luna — descending in capability and price, currently available only to U.S. government-approved organizations via API and Codex. GPT-5.6 Sol, the flagship tier, features a new 'max reasoning' mode and 'ultra mode' that spawns multiple subagents for multi-step tasks, and achieved state-of-the-art results on Terminal-Bench 2.1 (91.9%) while approaching Claude Mythos 5 on ExploitBench. The models include layered biosecurity and cybersecurity guardrails, with independent evaluations from METR and SecureBio yielding mixed but notable findings — particularly a near-10-point biology knowledge jump over GPT-5.5 and ambiguous autonomous task-duration results from METR. Wider public release is planned within weeks.