Microsoft ships its first streaming transcription model, built for real-time voice agents
Microsoft completes the real-time transcription piece of its voice-agent stack, but the latency and pricing figures are the company's own and await developer validation.
Microsoft released MAI-Transcribe-2-Streaming, the first streaming transcription model in its MAI family, on October 1. It keeps producing provisional transcripts while the speaker is still talking.
According to Microsoft, the model accepts speech over a WebSocket, returns its first transcript hypotheses within 320 milliseconds on average, supports more than 60 languages with automatic language detection, and is priced at 54 cents per audio hour on the Vercel AI Gateway. The company itself notes that real-world speed also depends on the network connection and the model generating the reply, so the latency figure is not guaranteed in every scenario.
For comparison, the non-streaming MAI-Transcribe-2 released last month costs 10 cents per audio hour and waits until the speaker finishes before processing; the streaming variant costs more than five times as much in exchange for live output. Microsoft also shipped two text-to-speech models the same day, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Combined with its Mai-Thinking-1 reasoning model, Microsoft now has the components for a full voice agent built in-house, consistent with its stated goal of reducing what it pays OpenAI and Anthropic.
Sources:https://microsoft.ai/news/our-first-streaming-transcription-model