Microsoft launches real-time streaming transcription model, tops third-party leaderboard
Microsoft's first streaming transcription model outputs text while speech is ongoing, ranking first among 28 models; production performance remains to be seen.
Microsoft announced on October 1 its first real-time streaming speech transcription model, MAI-Transcribe-2-Streaming, which outputs text continuously while speech is ongoing, covering 60 languages with automatic language detection.
The model is in a promotional pricing phase at $0.54 per hour, about $9 per 1,000 minutes; the earlier non-streaming MAI-Transcribe-2 was priced at $0.10 per hour in Azure Speech public preview. Developers can access it through Microsoft Foundry, MAI Playground and OpenRouter.
Per Microsoft's announcement, Artificial Analysis's September 28 streaming transcription leaderboard measured the model at a 2.50% final word error rate and 0.13-second final latency, both first among 28 models, ahead of Grok Voice Transcribe 2.0 Streaming at 2.73%. Voice applications can now start understanding or calling tools before a speaker finishes a sentence; real-world production performance remains to be verified.