sakana-ai-acc49fdb·3 events·first seen Aliases: Sakana AI
Sakana AI, a Tokyo-based research lab, released two dedicated orchestrator models—Fugu and Fugu-Ultra—that dynamically delegate tasks to a pool of underlying LLMs including Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5 under a single API. Fugu-Ultra achieves state-of-the-art results on SWE-Bench Pro, Humanity's Last Exam, LiveCodeBench Pro, and GPQA-Diamond, outperforming individual frontier models on several benchmarks. The models are trained via supervised fine-tuning plus sep-CMA-ES evolutionary optimization and GRPO reinforcement learning to select the best worker model per subtask, with Fugu-Ultra using a sub-component called Conductor to coordinate parallel agentic workflows. The approach represents a commercially available alternative to dependence on any single frontier model, with pricing available via Sakana API, OpenRouter, and Vercel.
OpenAI and Broadcom announced Jalapeño, OpenAI's first custom inference chip, designed in nine months with AI-assisted design and showing better performance-per-watt than current accelerators; engineering samples are already running GPT-5.3-Codex-Spark with datacenter deployment planned by end of 2026. Sakana AI released Fugu, a multi-agent routing system that scored 73.7% on SWE-Bench Pro, outperforming Claude Opus 4.8 and GPT-5.5 while remaining below the inaccessible Fable 5. Additional items cover Anthropic's Claude Tag Slack integration for async team collaboration, Seedance 2.5 video model improvements, the Robin autonomous biology research agent that identified a novel drug candidate, and a Getty Images licensing partnership with OpenAI.
The Batch analyzes the surge of interest in recursive self-improvement (RSI) triggered by Anthropic's report that Claude now authors or co-authors 80% of the company's code, up from under 5% before Claude Code launched. The piece documents concrete productivity metrics—engineers contributing 8x more code lines in Q2 2026 versus Q1 2023, and 800 API fixes shipped in April that would have taken humans four years alone—alongside a spectrum of community reactions ranging from skeptical (Brundage, Mollick) to opportunistic (OpenAI, Sakana AI's new RSI Lab). The commentary frames RSI as theoretically distant but notes the marketing dimension of Anthropic's framing and the gap between agentic coding assistance and true self-directed improvement.