
claude-sonnet-3-5-637fb510·10 events·first seen Aliases: Claude Sonnet 3.5, Claude Sonnet 5
OmniaBench is a new benchmark for evaluating general AI agents across diverse scenarios, spanning 90 level-1 and 354 level-2 domains derived from app stores, product documents, and web retrieval. The benchmark contains 1,431 tasks with single-turn and multi-turn formats, a ten-dimensional capability taxonomy, and eight atomic difficulty factors for fine-grained analysis. Even frontier models like Claude Sonnet 5 and GPT-5.6-Sol achieve only ~58% Overall Pass@1, revealing persistent weaknesses in planning, constraint maintenance, and adaptive correction. The work addresses a gap in existing agent benchmarks that tend to cover narrow tool ecosystems or interaction formats.
RuBench 1.0 is a new benchmark of 25 repository-level agentic coding tasks drawn from real fix commits in five live open-source projects, where task specifications are written natively in Russian in the style of customer requests rather than translated from English. The benchmark evaluates deployed product configurations including Claude Code with Opus 4.8, Sonnet 5, and Haiku 4.5, and Codex CLI with GPT-5.5, with the best configuration resolving 78.7% of tasks. A notable finding is that auditing trajectories of a fifth configuration (Claude Code + Fable 5) revealed that on 20% of tasks an official safeguard fallback silently re-routed the model to Opus 4.8, providing direct evidence that the deployed product rather than the underlying model is the actual unit of measurement in agentic evaluations.
Claude Code release 2.1.201 removes the use of the mid-conversation system role for harness reminders in Claude Sonnet 5 sessions. This is a narrow behavioral change to how the Claude Code agent harness communicates with the model during sessions. The change likely reflects improvements in Sonnet 5's ability to maintain context without mid-session system-role injections.
The Trump administration lifted export restrictions on Anthropic's Claude Fable 5 and Claude Mythos 5 after Anthropic committed to stronger safeguards, resolving a dispute over jailbreak vulnerabilities. Separately, Anthropic launched Claude Sonnet 5, a mid-tier agentic model priced at $2/$10 per million tokens through August 2026, and Claude Science, a unified research workbench for life sciences integrating PubMed, Jupyter, and HPC cluster access. The newsletter also covers Google's Nano Banana 2 Lite image model and Gemini Omni Flash video model, and Cognition's Devin Fusion multi-model routing system claiming 35% cost reduction versus GPT-5.5 and Opus 4.8.
Zvi Mowshowitz (Don't Worry About the Vase) publishes commentary on Claude Sonnet 5, characterizing it as a non-frontier model with specific practical applications. The post is gated for premium subscribers after a one-week free window. The piece appears to be part of Zvi's ongoing 'Fable' series of AI model evaluations.
Latent Space's AINews digest covers the release of Claude Sonnet 5 and previews Fable 5, suggesting both are significant near-term developments in the AI landscape. The newsletter aggregates community and industry signals around these releases. The brief body ('Everything is open again!') suggests a theme around open-weights or open-access model availability.
Anthropic has released Claude Sonnet 5, now the default model in Claude Code, featuring a native 1M-token context window. The model is available at promotional pricing of $2/$10 per million tokens (input/output) through August 31, 2026. Users must update to Claude Code version 2.1.197 to access the new model.
Anthropic has released Claude Sonnet 5, a new mid-tier model in their Claude lineup. The announcement comes via the official Anthropic news page and generated significant community engagement on Hacker News with 714 points and 386 comments. As a new named model release from a frontier lab, this is a notable update to the Claude model family.
Simon Willison published a commentary post on Claude Sonnet 5, covering what is new in the release. The post is from a tier-2 source and represents secondary analysis of Anthropic's model update. The body content was not provided, so specific claims cannot be assessed, but the subject is a notable mid-tier model release from Anthropic.
Anthropic released a post-mortem on AI and elections in 2024, covering their safety policies, red-teaming efforts, and enforcement actions across global elections. Election-related activity constituted less than 0.5% of overall Claude usage, rising to just over 1% around the US election, with approximately 100 enforcement actions globally. The report introduces Clio, an automated tool for analyzing real-world usage patterns, and documents a case study on handling knowledge cutoff limitations during France's snap elections. The piece represents Anthropic's first systematic public accounting of election-related AI safety work at scale.