By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Microsoft AI Releases MAI-Transcribe-2-Streaming Speech-to-Text Model
Microsoft AI launched MAI-Transcribe-2-Streaming on October 1, 2026, marking its inaugural streaming speech-to-text (STT) model. This release coincided with two new text-to-speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Independent analysis by Artificial Analysis has positioned MAI-Transcribe-2-Streaming as the top-performing model out of 38 evaluated for both final and initial partial transcript accuracy. The model is specifically engineered for applications where low latency is critical, such as voice agents, live captioning, and dictation software. MAI-Transcribe-2-Streaming functions as the real-time counterpart to MAI-Transcribe-2, a batch processing STT model introduced in September. It supports transcription for 60 languages and features automatic, continuous language detection. The system processes audio streams in real-time, delivering text output while the speaker is still actively talking. It generates its first text hypotheses, referred to as partials, within approximately 100 milliseconds of receiving audio input. As more context becomes available, these partial transcripts are refined before a stable, final transcript is committed. This capability allows AI agents to begin processing information or executing commands even before a sentence is fully completed. Microsoft's internal testing indicates that words appear in the transcript up to twice as fast as its nearest competitor.
Artificial Analysis employed the AA-WER Streaming index for its evaluation, utilizing approximately 8 hours of audio data. This dataset comprised a mix of audio sources: 50% from AA-AgentTalk, 25% from VoxPopuli, and 25% from Earnings22. Latency measurements were initiated from the point where speech concluded, as identified by the SileroVAD algorithm. The model achieved a Word Error Rate (WER) of 2.5% at 0.13 seconds after the end of speech, securing the #1 ranking among 38 models. For first partial transcripts, it also achieved a 2.5% WER at 0.12 seconds after the end of speech, again ranking first. Competitors included Grok Voice Transcribe 2.0, which recorded a 2.7% WER at 0.49 seconds, and Muse Voice Transcribe, with a 3.1% WER at 0.16 seconds. While Cartesia Ink-2 (using external endpoints) provided final transcripts in a faster 0.07 seconds, it did so with a higher WER of 4.0%. The benchmark highlights that the initial partial transcript offers accuracy comparable to the final transcript, a crucial factor for AI agents that need to act preemptively. Microsoft has also positioned MAI-Transcribe-2-Streaming on the Pareto frontier, balancing accuracy against latency.
Microsoft AI is a division of Microsoft Corporation focused on developing artificial intelligence technologies. Artificial Analysis is an independent entity that benchmarks AI models. WER, or Word Error Rate, is a standard metric for evaluating the performance of speech recognition systems, measuring the number of substitutions, deletions, and insertions required to transform the recognized text into the reference text. Lower WER indicates higher accuracy. Latency in STT refers to the delay between spoken audio and the appearance of its transcribed text, a critical factor for real-time applications.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.