OpenAI launches GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper as the Realtime API exits beta
OpenAI launched three new real-time audio models on May 7, 2026 — GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper — while simultaneously moving the Realtime API from research preview to general availability, making it production-ready for the first time.
What's new
The three models cover distinct use cases:
- GPT-Realtime-2: A voice model with GPT-5-class reasoning, designed for voice interfaces that can handle complex requests and carry conversations forward. Billed by token consumption.
- GPT-Realtime-Translate: A live translation model supporting 70+ input languages and 13 output languages, translating speech in real time while keeping pace with the speaker. In evaluations across Hindi, Tamil, and Telugu, it delivered 12.5% lower Word Error Rates than competing models, with lower fallback rates and latency that sustained natural conversation. Billed by the minute.
- GPT-Realtime-Whisper: A streaming speech-to-text model that transcribes audio as people speak, built for low-latency captioning, meeting notes, and voice interfaces that need to feel immediately responsive. Also billed by the minute.
The Realtime API's GA status removes the research-preview designation, meaning developers can build production workloads against it with standard SLA expectations.
Conversations can be halted if they are detected as violating OpenAI's harmful content guidelines — a safety control that applies across all three models in real time.
Context
The original Realtime API launched in late 2024 and has been in beta or research preview since. The May 2026 release is the first time OpenAI has committed to it as a production-grade, generally available product, which is a prerequisite for enterprise adoption in customer service, healthcare, and financial services — industries that have been slow to adopt AI voice due to reliability requirements.
The GPT-Realtime-2 model's GPT-5-class reasoning is the key upgrade over prior versions: earlier Realtime API models were capable of conversational speech but could not handle requests requiring multi-step reasoning or task completion. The new model aims to close that gap.
GPT-Realtime-Translate directly targets the live multilingual market that has historically required human interpreters or post-session translation. The 70+ input language count is notably broader than competing models, and the 12.5% WER improvement in South Asian languages suggests specific training investment in those markets.
Why it matters
Voice is the interface modality that reaches the widest range of users, including those who are not fluent typists or who prefer hands-free interaction. Moving the Realtime API to GA opens it to industries that previously could not accept beta-grade risk.
The translation capability is potentially transformative for customer support and global commerce: a single agent model that can handle real-time translation across 70+ languages reduces the operational complexity of running multilingual call centers. The billing structure — by the minute for Translate and Whisper, by token for the reasoning model — allows cost-conscious buyers to optimize based on use case.
For competitors, the GA milestone raises the bar: Deepgram, AssemblyAI, and Google's speech-to-text services now face a well-capitalized counterpart that also ships reasoning capabilities alongside transcription.
Corroborating sources
- Openai
https://openai.com/index/advancing-voice-intelligence-with-new-models-in-the-api/
“Together, the models we are launching move real-time audio from simple call-and-response toward voice interfaces that can actually do work: listen, reason, translate, transcribe, and take action as a conversation unfolds”
- Techcrunch
https://techcrunch.com/2026/05/07/openai-launches-new-voice-intelligence-features-in-its-api/
“Together, the models we are launching move real-time audio from simple call-and-response toward voice interfaces that can actually do work”