By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Cartesia Sonic-3.6 Leads Speech Synthesis Benchmarks
Cartesia has released Sonic-3.6, the latest iteration of its real-time text-to-speech (TTS) model, approximately three months after the introduction of Sonic-3.5. The primary advancement in Sonic-3.6 is its enhanced naturalness, a quality that Cartesia asserts is independently verifiable. This new version now holds the number one position on both of Artificial Analysis's speech leaderboards. Specifically, it achieved an Elo rating of 1,283 on the Provider Voice board and 1,123 on the Controlled Voice board. The Controlled Voice board is considered more significant as it standardizes the evaluation by cloning every model onto the same eight reference voices, thereby isolating the synthesis engine's performance from the voice catalog. In this critical comparison, Sonic-3.6 leads, with Sonic-3.5 securing the second position and ElevenLabs' Eleven v3 ranking third.
Architecturally, Sonic-3.6 utilizes state space models instead of the more common transformer architecture. Cartesia claims that this architectural choice enables a time-to-first-audio latency of under 90 milliseconds. The model is currently available in beta and can be accessed via a hosted API, though it is not offered as self-hosted weights. Sonic is a closed, commercial model, meaning there are no open weights or a Hugging Face repository; users rent access to the service. Cartesia offers tiered pricing plans catering to various user segments, including solo developers and startups with Free and Pro ($5) tiers, scaleups managing contact centers with Startup ($49) and Scale ($299) plans, and regulated enterprises requiring Data Processing Agreements (DPAs), Business Associate Agreements (BAAs), and Single Sign-On (SSO) capabilities.
The applications for Sonic-3.6 span numerous industries and use cases. In financial services, healthcare, retail and e-commerce, logistics, recruiting, SaaS support, and consumer companion apps, the model can power inbound support agents, outbound qualification calls, and replace Interactive Voice Response (IVR) systems. It is also suitable for appointment reminders, sales-training simulators, audio localization for media, and providing voice user interfaces within products. Cartesia's launch page frames the inherent trade-offs in TTS—speed versus naturalness, and accuracy versus cost—as architectural challenges rather than insurmountable limitations. The practical implication of this approach is the stated low latency. Cartesia reports sub-90ms TTS latency for Sonic-3.6 and a 100ms transcript latency for its accompanying Ink-2 speech-to-text model. These figures represent vendor-stated model latencies, not necessarily measured end-to-end round-trip times.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.