
Sunday: when AI safety labs become the incident report
This week's throughline is trust — in evaluations, in model outputs, in the labs running both. From real-world unauthorized access during safety testing to frontier models fabricating biased medical diagnoses, the gap between what AI systems are supposed to do and what they actually do got harder to ignore.
Anthropic discloses three real-world unauthorized access incidents during cybersecurity evaluations
Real organizations were breached during safety evals — a reminder that evaluation infrastructure is itself an attack surface, and miscommunication between labs and partners has concrete consequences.
Frontier VLMs confabulate demographic-biased diagnoses when no medical image is provided
Frontier VLMs fabricating demographically biased diagnoses — while hedging in prose — shows that structured outputs can fail silently in ways prose-only audits will miss.
OpenAI announces ten advances in mathematics and theoretical computer science
Results across geometry, cryptography, and complexity theory in one announcement signals that AI-assisted mathematical research may be moving faster than the field expected.
OpenAI releases GPT-5.6 with improved efficiency across models, inference, and agentic workflows
GPT-5.6 reframes the frontier race around efficiency as much as raw capability — useful intelligence per dollar is now the stated metric.
Get the next edition in your inbox
Midweek and Sunday. Primary-sourced AI news paired with the guides that explain it.