big-bench-audio-27d42a2a·3 events·first seen Aliases: Big Bench Audio, BigBench Audio
X³-OPD is a new training framework that distills reasoning capabilities from a powerful text-based teacher model into an audio-language student model via on-policy alignment. The approach generates reasoning trajectories conditioned on the student's own acoustic perception while the teacher provides token-level guidance from matched textual inputs. A three-tier symmetric corpus covers speech-rendered text reasoning, audio-event reasoning, and paralinguistic spoken-dialogue reasoning. Evaluations on MMSU, MMAU, BIG Bench Audio, and MMAR show substantial improvements in audio-grounded reasoning and chain-of-thought quality.
Thinking Machines Lab (founded by Mira Murati) has announced TML-Interaction-Small, a 276B-parameter mixture-of-experts multimodal model that processes audio, video, and text concurrently using 200ms 'micro-turns' rather than waiting for conversational turns to complete. The architecture uses encoder-free early fusion, pairing a fast foreground interaction model with an asynchronous background reasoning model that shares context. On interactivity benchmarks (FD-bench V1/V1.5), it outperforms GPT-Realtime-2 and Gemini-3.1-flash-live-preview, though it trails GPT-Realtime-2 on intelligence benchmarks. A closed research preview is expected in coming months with wider release later in 2026.
Hugging Face introduces Big Bench Audio, a new benchmark designed to evaluate audio reasoning capabilities in AI models. The benchmark appears to extend the Big Bench evaluation framework into the audio domain, targeting multimodal models that process and reason over audio inputs. This release addresses a gap in evaluation tooling for audio-capable language models.