Interestana
Home/News/Superwhisper Releases Open-Weights Text Normalizer S1-mini
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Superwhisper Releases Open-Weights Text Normalizer S1-mini

Superwhisper has released the S1 family of models, comprising S1-Voice, S1-Language, and S1-mini. S1-Voice functions as a cloud-based speech-to-text model, while S1-Language is a cloud-based instruction-following model for text cleanup and formatting. The S1-mini model, however, stands out due to its release with open weights on Hugging Face, making it accessible for broader deployment. S1-mini is specifically a text normalizer, not an automatic speech recognition (ASR) transcriber or a chat model. Its primary function is to process raw transcripts generated by ASR systems and transform them into clean, readable written text. This process involves removing filler words, resolving self-corrections to reflect the speaker's final intent, applying correct punctuation and capitalization, and converting spoken numbers, dates, currency, and email addresses into their written forms.

The S1-mini model is fine-tuned from Qwen/Qwen3-0.6B and, in its initial release version 1, exclusively supports English. Its operation is controlled by a three-axis control line positioned above the input transcript. Superwhisper reports a token accuracy of 94.8% on a held-out dataset comprising 7,519 cases, with measurements taken greedily on the quantized build. This level of accuracy indicates its effectiveness in standardizing ASR output. The model's deployability is a key feature, with S1-mini being the self-hostable component of the S1 family. S1-Voice and S1-Language are offered as Superwhisper-hosted services, meaning they are consumable via API but not available for self-hosting.

S1-mini is published on Hugging Face under the Apache 2.0 license, with an additional clause regarding naming conventions. The Q4_K_M GGUF build of S1-mini is a compact 462 MB file, enabling it to run on a laptop's CPU. This small footprint allows solo developers to integrate it into desktop applications. For enterprises, S1-mini can be deployed within a Virtual Private Cloud (VPC), ensuring that audio transcripts remain within the organization's network, thereby enhancing data privacy and security. The model's capabilities are relevant across various industries, including healthcare and clinical documentation, legal services, financial services, customer support, developer tooling, and accessibility applications like live captioning. Its applications extend to dictation software, meeting note tools, live captioning services, voice-driven editors, and systems for direct voice-to-CRM entry, as well as any pipeline that converts raw ASR output into human-readable text.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next