Interestana
Home/News/Alibaba Qwen Releases Real-Time Translation Model
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Alibaba Qwen Releases Real-Time Translation Model

Alibaba's Qwen team has released Qwen3.8-LiveTranslate, a new generation of its real-time simultaneous interpretation model. This model is designed to process live speech, with the option to include video frames, and deliver translated text and speech concurrently with the original speaker. The primary innovation in Qwen3.8-LiveTranslate is its novel Interleave architecture, which the Qwen team reports enhances faithfulness, fluency, and conciseness in translations. A key performance metric, Length-Adaptive Average Lagging (LAAL), has been reduced from an average of 2.8 seconds to 2.3 seconds, representing an approximately 18% decrease in translation delay. This improvement allows the translated output to trail the source speech by a shorter duration on average, while also discouraging excessive output generation. The model is available as a hosted API and can be accessed through Alibaba Cloud Model Studio and QwenCloud under the identifier qwen3.8-livetranslate-flash-realtime via WebSocket. The Qwen3.8-LiveTranslate model is described as the real-time counterpart to the Qwen3.8-LiveTranslate-Flash model, which supports offline audio and video translation. It is built upon the Qwen-Omni stack and leverages large-scale multimodal data, cross-language and cross-modal alignment techniques, and visual enhancements. Beyond the core translation improvements, Qwen3.8-LiveTranslate introduces three significant new capabilities. First, it offers real-time speaker diarization, enabling the model to distinguish between different speakers in multi-party conversations. This feature also aims to maintain speaker identity through more stable voice cloning. The API provides different voice cloning modes, including an 'always' mode that re-clones a speaker's voice before each response in multi-speaker scenarios to ensure consistent voice cloning. Second, the model supports synchronized bilingual display, where the original source text and its translation appear on screen simultaneously. Within the API, the source transcription is streamed as separate events alongside the translation stream. Third, it incorporates long-context disambiguation, a feature designed to improve translation accuracy by considering a broader scope of the conversation or text to resolve ambiguities. This capability is crucial for maintaining coherence and precision in longer discussions or documents. The Qwen team's focus on reducing latency while maintaining translation quality addresses a critical challenge in simultaneous interpretation, making the technology more practical for live communication scenarios across a wide range of languages.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next