OpenAI released two new transcription models: GPT Transcribe for accurate file transcription and final transcripts of committed Realtime API turns, and GPT Live Transcribe for low-latency streaming transcription. Both models support free-form transcription context, keyword hints, and multiple expected input languages. The release extends OpenAI's speech-to-text capabilities with a dedicated streaming path alongside a high-accuracy batch path.
OpenAI introduced three new audio models in its Realtime API: GPT-Realtime-2 (speech-to-speech with five configurable reasoning effort levels), GPT-Realtime-Translate (70+ input languages), and GPT-Realtime-Whisper (transcription). GPT-Realtime-2 operates as an end-to-end audio model including reasoning, with latency ranging from 1.12 seconds at minimal effort to 2.33 seconds at high effort. Benchmark results are mixed: it leads Scale AI's Audio MultiChallenge and Artificial Analysis Conversational Dynamics but trails Step-Audio R1.1 Realtime and Grok Voice Think Fast 1.0 on speech reasoning and agentic tasks. The configurable reasoning-latency tradeoff is positioned as a key differentiator for voice agent applications.
OpenAI is deploying a new behind-the-scenes speech-to-text model for dictation in ChatGPT across all plans. The update improves transcription accuracy across multiple languages and accents, including multilingual code-switching, noisy environments, and whispered speech. Internal evaluations show at least 10% reduction in word error rate for top languages compared to the previous production model.
OpenAI released GPT-Realtime-2.1, an updated realtime reasoning model with improvements to alphanumeric recognition, silence and noise handling, and interruption behavior. A companion model, GPT-Realtime-2.1 mini, was also released as a faster, lower-cost distilled variant for realtime voice use cases. The releases represent incremental improvements to OpenAI's realtime voice API tier rather than a flagship capability shift.
OpenAI announced GPT-Live, a new generation of voice models designed for natural human-AI interaction, now powering ChatGPT Voice. The announcement comes from OpenAI's official blog, indicating a production-grade voice capability release. This represents a significant update to OpenAI's real-time voice interaction stack.
OpenAI has released a suite of new real-time voice and audio APIs including GPT-Realtime-2, a GPT-Translate model, and an updated Whisper, all positioned as state-of-the-art for real-time voice applications. The releases appear to be part of a broader push to deploy GPT-5 capabilities across multiple product surfaces. Coverage comes from the Latent Space AI News digest, which aggregates and contextualizes the announcements.
OpenAI has updated the floating model slugs for gpt-4o-mini-tts and gpt-4o-mini-transcribe to point to their 2025-12-15 snapshots, with the previous March 2025 snapshots remaining accessible via versioned identifiers. Notably, OpenAI now recommends gpt-4o-mini-transcribe over gpt-4o-transcribe for best transcription results, signaling a quality improvement in the mini-tier audio model.
OpenAI has released gpt-realtime-1.5 to the Realtime API and gpt-audio-1.5 to the Chat Completions API. These are incremental model updates to OpenAI's audio and real-time speech capabilities. The release expands developer access to updated audio-capable models through existing API surfaces.
OpenAI is rolling out GPT-Live-1 to power ChatGPT Voice for paid users, with GPT-Live-1 mini for Free users. Both models support simultaneous listening and speaking, enabling more natural turn-taking and interruptions. GPT-Live-1 integrates with web search, memory, visual widgets, and multimodal (text and image) input within a single chat session, though video and screen sharing remain exclusive to the existing Advanced Voice Mode. The rollout targets consumer plans on chatgpt.com and mobile apps, excluding Business, Enterprise, and Edu workspaces at launch.