Interestana
Home/News/Alibaba Qwen Releases Omni-Modal AI Model With 1M Context
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Alibaba Qwen Releases Omni-Modal AI Model With 1M Context

Alibaba Qwen Releases Omni-Modal AI Model With 1M Context

Alibaba's Qwen team has released Qwen3.8-Omni-Flash, their first omni-modal artificial intelligence model designed with agentic capabilities. This new model is capable of processing multiple types of data, including text, images, audio, and video, and it generates text outputs. A key feature of Qwen3.8-Omni-Flash is its integrated approach to audio-video understanding, reasoning, and tool utilization within a single model architecture. The operational workflow is described as a straightforward process: first, understand the input content, then plan the required task, subsequently execute the task using available tools, and finally deliver the desired result. Qwen3.8-Omni-Flash is currently available as a hosted API through Alibaba Cloud's QwenCloud, Model Studio, and Qwen Studio. At the time of its launch, no open weights were announced, meaning self-hosting is not an option for users. The Qwen3.8-Omni-Flash model is built upon the Qwen3.8-Flash-Next architecture, a base model for which open weights were released in August 2026. The model boasts an extensive context window of 1 million tokens, with QwenCloud specifying a maximum input of 991,000 tokens and a maximum output of 131,000 tokens. The maximum reasoning length is stated to be 262,000 tokens, with all outputs being text-only. Developers requiring generated speech capabilities are directed to use Qwen3.5-Omni. The model's "thinking" capability is enabled by default, with the `reasoning_effort` parameter set to `xhigh`; disabling this feature is possible by setting it to `none`. The API is designed to be compatible with both the DashScope and OpenAI protocols, supporting functionalities such as Chat Completions and the Responses API, along with features like function calling, web search integration, structured outputs, context caching, and batch processing. A significant advancement highlighted by the Qwen research team is the "Agentic Perception for Long Video" capability. Unlike conventional video models that process entire files sequentially, even when the relevant information is localized, this agentic approach begins with the user's question. It then intelligently decides which segments of the video and audio to process, gathering evidence through multiple rounds of coarse-to-fine analysis. This method directs computational resources and tokens primarily to the most pertinent segments of the media. The research team reported improvements on the OmniVideoBench benchmark, showing an accuracy increase from 63.4% to 67.8%, while simultaneously reducing token usage by approximately 45.7%, from 145,736 tokens to 79,117 tokens. All benchmark figures provided are attributed to Qwen, as independent verification was not available at the time of publication. Across 29 different evaluations, the average score for Qwen3.8-Omni-Flash demonstrated an improvement exceeding 25% compared to previous Qwen models.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next