By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Alibaba Tongyi Lab Releases Qwen-Audio-3.0-TTS in 16 Languages
Alibaba's Tongyi Lab released Qwen-Audio-3.0-TTS, a production-ready text-to-speech system available in two tiers: Flash for real-time interaction and Plus for high-quality generation. Both variants are hosted on Alibaba Cloud Model Studio and support 16 languages, addressing developer needs for broader language coverage, natural-language style control, fine-grained tag control, and robustness with noisy reference audio. The Qwen-Audio-3.0-TTS-Plus model has achieved the top rank on the Artificial Analysis Text-to-Speech leaderboard.
The Flash tier is optimized for low latency, with first-packet delivery at approximately 300 milliseconds, making it suitable for interactive applications. The Plus tier prioritizes naturalness and timbre fidelity for applications where audio quality is paramount. Both models, identified as qwen-audio-3.0-tts-flash and qwen-audio-3.0-tts-plus, are accessible via a bidirectional WebSocket streaming protocol. The API supports various audio formats including PCM, WAV, MP3, and Opus, with output sample rates up to 48 kHz. Key features include streaming input and output, voice cloning, Voice Design capabilities, and instruction control.
Alibaba provides the DashScope SDK and example code in multiple programming languages such as Python, Java, Go, C#, PHP, and Node.js. These resources are available for deployment in Alibaba's Singapore and Beijing cloud regions. The underlying architecture of Qwen-Audio-3.0-TTS incorporates a 12.5 Hz low-frame-rate speech tokenizer, which reduces autoregressive decoding costs by decreasing the number of tokens per second of audio, thereby lowering inference latency. This design choice helps retain essential content and speaker information while improving efficiency.
The model's development involved a five-stage progressive training paradigm. This pipeline integrates language model (LM) and flow-matching (FM) components through independent pretraining, joint training with high-quality data annealing, LM reinforcement learning, FM robustness training, and FM reinforcement learning. This structured approach is reported by the research team to enhance the model's content generation capabilities and overall performance.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.