Interestana
Home/News/Alibaba Qwen Releases Realtime Full-Duplex Voice Model
MarkTechPost••4 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Alibaba Qwen Releases Realtime Full-Duplex Voice Model

Alibaba's Qwen team has unveiled Qwen-Audio-3.1, an advanced audio processing stack comprising five distinct models designed for artificial intelligence agents. The centerpiece of this release is Qwen-Audio-3.1-Realtime, a full-duplex speech model engineered to facilitate seamless, real-time interaction for voice-enabled AI systems that require tool-calling capabilities. Concurrently, Qwen has implemented substantial price reductions across its audio offerings, with decreases of approximately 85% for the Realtime model, around 70% for its Text-to-Speech (TTS) capabilities, and up to 95% for its Automatic Speech Recognition (ASR) services. This new model is accessible as a managed API, with Qwen-Audio-3.1-Realtime-Plus now available on QwenCloud via WebSocket. No open-weight versions of the model have been announced.

The QwenCloud deployment of Qwen-Audio-3.1-Realtime supports both text and audio as input and output modalities. It boasts a significant context window of 262,000 tokens, with a maximum input capacity of 245,000 tokens and a maximum output capacity of 16,000 tokens. The service operates under default limits of 60 requests and 100,000 tokens per minute. Pricing for the model on QwenCloud is set at $6.4 per 1 million audio input tokens and $0.8 per 1 million text input tokens. For output, both text and audio are priced at $24 per 1 million tokens, with text output not incurring additional charges beyond the token cost. Key functionalities integrated into the model include function calling, web search integration, generation of structured outputs, a context cache for efficient memory management, and fine-tuning capabilities.

Complementing the Realtime model, Alibaba has also introduced Qwen-Audio-3.1-ASR-Flash-Filetrans, a model specifically designed for offline transcription of long audio files. This ASR model features advanced capabilities such as hot word detection, speaker diarization (separation), automatic punctuation insertion, and support for multilingual recognition, including various Chinese dialects. The pricing for Qwen-Audio-3.1-ASR-Flash-Filetrans is $0.15 per 1 million input tokens and $0.47 per 1 million output tokens.

The underlying architecture of the Qwen-Audio-3.1 system involves two primary models that share a common Audio Encoder and Large Language Model (LLM) design. A full-duplex decision-making model is responsible for predicting the optimal action at any given moment: whether to continue listening, initiate speech, cease speaking, or resume listening. This is complemented by a speech-to-text model that transcribes the spoken content into text. Subsequently, a context-aware voice renderer converts this text into a streaming speech output, taking into account conversation history, vocal cues, and acoustic context to produce natural-sounding speech. The training methodology for this system is structured into three distinct layers: Think, Act, and Speak and Coordinate. The 'Think' layer utilizes M²-OPD Core-Cocktail SFT, a process that re-anchors the audio model to its source text LLM using extensive paired data, reportedly on the scale of millions of hours. This is followed by Multimodality OPD. In this stage, a Text Teacher model and a frozen Audio Reference model collaboratively score each token of the student model's generated trajectory. This approach is described as on-policy distillation, focusing on learning from the model's own actions rather than imitating pre-written responses. The training also incorporates domain experts to imbue the model with qualities like empathy and pragmatic understanding.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next