A security researcher reports discovering a WordPress remote code execution vulnerability using GPT-5.6 at a cost of approximately $25 in API spend, contrasting with the $500,000 market price exploit brokers pay for such findings. The post demonstrates a concrete offensive security capability enabled by frontier LLMs. This is a significant capability signal showing that high-value vulnerability research is becoming accessible at near-zero cost with current models.
OpenAI has launched a red-teaming bug bounty program specifically targeting biosafety risks in GPT-5.5, offering rewards up to $25,000. The program focuses on finding universal jailbreaks that could bypass biological safety guardrails. This represents a structured external adversarial evaluation of a frontier model's safety properties in a high-stakes domain.
A developer built a deliberately vulnerable application and ran LLMs against it as automated penetration testers, spending $1,500 on API costs across the experiment. The post evaluates how well current LLMs can identify and exploit real vulnerabilities in a controlled setting. Results provide practical signal on the current state of LLM-assisted offensive security, a capability area with both red-team and safety implications.
OpenAI announced a Bio Bug Bounty program specifically targeting GPT-5.5, inviting external researchers to probe the model for biosecurity-related vulnerabilities or capability failures. The program is a structured red-teaming and safety evaluation initiative focused on biological risk. This signals continued investment in pre- and post-deployment safety evaluation for frontier models, particularly in high-stakes domains.
A Google security report catalogs emerging LLM-enabled cyberattack techniques including morphing malware with mutation engines, logical-flaw discovery in code, and AI-directed obfuscation networks. The report was prompted in part by a real incident where hackers used an LLM to find a zero-day in a widely used web administration tool. Separately, the UK AI Security Institute found that Claude Mythos Preview and GPT-5.5 can reliably execute attacks expected to take humans 3 hours, up from earlier 1-hour benchmarks, with performance scaling further when token limits are relaxed. The findings suggest an accelerating gap between LLM offensive capability and conventional defensive tooling.
OpenAI released GPT-5.4 in Thinking and Pro variants, featuring an expanded context window (up to 1.05M input tokens), native computer use, tool search capabilities, and adjustable reasoning levels. In independent testing by Artificial Analysis, GPT-5.4 Pro at xhigh reasoning achieved state-of-the-art on GDP-Val-AA, BrowseComp, Terminal-Bench-Hard, SWE-Bench-Pro, and MCP Atlas, while trailing Gemini 3.1 Pro Preview on MMMU-Pro and Humanity's Last Exam. Pricing is set at the top of the market ($30/$180 per million input/output tokens for Pro), and the release also powers Codex, OpenAI's competitor to Claude Code. The item is reported via The Batch (tier 2 commentary) and includes additional context on Andrew Ng's chub CLI tool for agent documentation sharing.
OpenAI published a blog post describing how GPT-5 is being used for medical research applications. The post appears to be an announcement or case study highlighting GPT-5's capabilities in a healthcare/research context. Specific details about methods, benchmarks, or outcomes are not provided in the available text.
Researchers introduce VEXAIoT, an autonomous multi-agent framework using LLM-based reasoning to discover and exploit vulnerabilities in IoT environments. The system pairs a vulnerability detection agent with an attack execution agent, evaluated across 260 attack executions in IoTGoat and Metasploitable2 environments covering ten OWASP IoT vulnerability categories. It achieves a 95% overall success rate with average execution times under two minutes, demonstrating that LLM agents can automate offensive IoT security workflows at scale in controlled settings.
OpenAI launched a preview of three vision-language models — GPT-5.6 Sol, Terra, and Luna — descending in capability and price, currently restricted to U.S. government-approved organizations. GPT-5.6 Sol is positioned as comparable to Claude 5 Mythos and claims state-of-the-art on Terminal-Bench 2.1; it includes a 'max reasoning' mode and an 'ultra mode' that delegates work to multiple agents. Pricing ranges from $5/$30 per million input/output tokens for Sol down to $1/$6 for Luna, with wider public access promised within weeks. All models include safeguards against dangerous biological, chemical, and cybersecurity information, with relaxed-safeguard variants also available to approved partners.