By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Gradium AI Releases New TTS Model With 81% Hard-Case Pass Rate
Gradium AI has released a new text-to-speech (TTS) model, making it the default for its API and Studio services as of August 31, 2026. This new model reportedly achieves an 81.0% human-rated pass rate on a 500-sentence "hard-case" evaluation set, designed to test its ability to accurately pronounce critical information often found in voice agent interactions. This set spans five languages: English, German, French, Spanish, and Portuguese. The company states this performance surpasses competitors, including Cartesia Sonic 3.6, which achieved a 75.1% pass rate, and ElevenLabs v3 Conversational, with a 65.4% pass rate. The evaluation set includes 100 sentences across 10 criteria, with seven atomic criteria focusing on spelling, acronyms, alphanumeric tokens, dates, regular numbers, large and floating numbers, and email addresses. Three composite criteria simulate realistic agent turns by combining multiple elements into orders, IT tickets, and claims. Scoring is rigorous, requiring independent native-speaker raters to confirm that every element of a sentence is pronounced correctly and completely; a single missed digit results in a sentence failure. The audio was loudness-normalized, and raters were limited to 40 comparisons with mandatory breaks to ensure consistent evaluation. The new model also demonstrates improved latency, with a P50 time-to-first-audio of 216 milliseconds on Coval’s TTS benchmark, which is 170 milliseconds faster than the model it replaces. This represents a significant improvement for applications requiring rapid voice responses, such as interactive voice response (IVR) systems and virtual assistants. Gradium AI has made the model available for immediate deployment without requiring migration for existing users. Custom voice clones and existing integrations will continue to function unchanged with the new default model. The company has open-sourced its 500-sentence evaluation set on Hugging Face under a CC BY 4.0 license, allowing other researchers and developers to benchmark their own TTS models. Other TTS models evaluated in August 2026 with default settings included Fish Audio S2.1 Pro at 49.5% and Inworld TTS 1.5 Max at 46.5%. The focus on "hard-case" scenarios highlights the ongoing challenge in voice AI to accurately handle the precise details that are crucial for user experience and task completion in customer service and other voice-driven applications. The ability to correctly pronounce order numbers, callback digits, and email addresses is vital for the functionality of voice agents, and Gradium AI's new model aims to address these specific failure points.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.