Home/News/Open ASR Models Compete on WER, Languages, Latency, and License
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Open ASR Models Compete on WER, Languages, Latency, and License

The landscape of open Automatic Speech Recognition (ASR) models has evolved beyond a single dominant player, with several models now competing on performance metrics and licensing. In March 2026, Cohere released Transcribe, a 2B parameter model under the Apache 2.0 license, which initially led the Hugging Face Open ASR Leaderboard with an average Word Error Rate (WER) of 5.42%. This was quickly followed by IBM's Granite Speech 4.1 2B, which achieved a WER of 5.33%. More recently, ARK-ASR-3B and MOSS-Transcribe-preview-2B have posted even lower WER scores, narrowing the gap at the top of the leaderboard to less than one WER point.

This intense competition means that rank alone is no longer the primary factor for users selecting an ASR model. Instead, critical considerations now include the model's license, the breadth of languages it supports, its streaming capabilities for real-time processing, and the cost per audio hour. These factors are becoming increasingly important for developers and businesses integrating ASR technology into their applications.

However, a closer examination of the Open ASR Leaderboard reveals nuances in how the reported average WER is calculated. The average WER cited is not a fixed quantity, and models have been scored using different test sets. For instance, Cohere's 5.42% WER is an average across eight English test sets, including TED-LIUM. In contrast, ARK-ASR-3B's 5.04% is an average across seven sets, excluding TED-LIUM. The MOSS-Transcribe-preview-2B card explicitly states that TED-LIUM is not part of its current leaderboard run. Since TED-LIUM is considered one of the easier test sets, its exclusion can artificially inflate the reported average WER.

When Cohere's published per-dataset scores are recomputed using the same seven test sets that ARK reports, Cohere's average WER increases to 5.84% from the headline 5.42%. Similarly, IBM's Granite Speech 4.1 2B moves from 5.33% to 5.65% when evaluated on the same basis. This suggests that, on a like-for-like comparison, ARK's lead over other models might be larger than initially indicated by the headline leaderboard figures. The focus is shifting towards transparent and consistent evaluation methodologies to accurately compare these rapidly advancing open ASR models.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next