Microsoft launches MAI-Transcribe-2, undercutting OpenAI, Google, and ElevenLabs on price and speed
Microsoft AI shipped MAI-Transcribe-2 on September 3, 2026, a new speech-recognition model the company says is the fastest, most accurate, and cheapest transcription model on the market, priced at $0.10 per hour of audio processed.
What's new
MAI-Transcribe-2 transcribes audio across 60 languages and adds speaker diarization, word-level timestamps, keyword biasing, and code-switching support for speakers who mix languages mid-sentence. VentureBeat reports that Microsoft "priced the thing at 10 cents per hour of audio," and that "Thursday's early-bird price cuts that by roughly 72%" from the prior model's rate, with Microsoft's launch throwing in "six features its rivals sell separately."
On throughput, Microsoft's own announcement claims MAI-Transcribe-2 runs roughly 10 times faster than OpenAI's GPT-Transcribe, seven times faster than ElevenLabs' Scribe v2, and five times faster than Google's Gemini 3.5 Transcribe. On accuracy, the model posts a 5.2% average word-error rate across 60 languages on the FLEURS benchmark, which Microsoft says puts it first on that leaderboard, ahead of Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and Scribe v2.
Context
Speech-to-text has become one of the more commoditized corners of the model market over the past year, with OpenAI, Google, and ElevenLabs all fielding transcription models aimed at developers building call-center, captioning, and meeting-notes products. Pricing in that segment has been falling steadily as providers compete on cost per hour of audio processed rather than on raw capability alone, since accuracy across the major offerings has converged for most mainstream languages.
Microsoft's entry follows that pattern: rather than leading with a new capability, MAI-Transcribe-2's pitch rests on undercutting the field on price while matching or beating it on speed and accuracy. The roughly 72% price cut from the prior model suggests Microsoft is treating transcription as a volume, infrastructure-margin business rather than a premium product line.
Why it matters
For developers building voice products, MAI-Transcribe-2 lowers the cost floor for transcription at scale, which matters most for applications that process large volumes of audio, such as call-center analytics, video captioning pipelines, and meeting-transcription tools, where per-hour pricing compounds quickly. A credible fivefold-to-tenfold speed claim, if it holds up under independent benchmarking, would also cut latency-sensitive use cases like live captioning.
More broadly, the launch signals that Microsoft is willing to compete directly on price in a category it doesn't need to own outright, since Azure already resells or integrates rival transcription APIs for many customers. Undercutting OpenAI, Google, and ElevenLabs on their own turf puts pressure on those companies to respond with price cuts of their own, which would be a net win for developers regardless of which vendor they choose.
Corroborating sources
- Venturebeat
https://venturebeat.com/infrastructure/microsoft-ais-mai-transcribe-2-undercuts-openai-google-and-elevenlabs-on-price-and-speed
“Then it priced the thing at 10 cents per hour of audio.”