Interestana
Home/News/Alibaba Qwen Releases 2.4T Parameter MoE Model
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Alibaba Qwen Releases 2.4T Parameter MoE Model

Alibaba Qwen Releases 2.4T Parameter MoE Model

Alibaba's Qwen team has made its Qwen3.8-Max model broadly available, with open weights scheduled for release the following week. This new model represents a significant advancement for the Qwen family, boasting a total of 2.4 trillion parameters and operating as a Mixture-of-Experts (MoE) architecture. Qwen3.8-Max is designed to process a multimodal range of inputs, accepting text, images, and video, and producing text-based outputs. The deployability of the model varies depending on the specific artifact used. For immediate use by companies of any size, a hosted API is available today. This API is compatible with OpenAI and DashScope, allowing for integration through a simple change of the base URL and model ID. The open-weights version of the model, however, is a different proposition. With its 2.4 trillion total parameters, the checkpoint is intended for multi-node datacenter environments. Alibaba has not yet disclosed the number of activated parameters, which means the cost of serving the model cannot be precisely modeled at this time. A second checkpoint, Qwen3.8-27B, is also slated for open-weights release and is designed to be compatible with standard on-premise GPU hardware, making it more accessible for individual deployment. The model's feature set is mapped to four key industries: software engineering, legal and financial document review, media and e-commerce operations, and design. Specific applications include the development of repository-scale coding agents and the creation of long-document knowledge bases. The model is also capable of long-video indexing, structured data extraction, and acting as a multi-step research assistant. The Qwen3.8-Max model features a substantial 1 million token context window. The maximum input capacity is 991,000 tokens, which slightly reduces to 983,000 tokens when the 'thinking' feature is enabled. The maximum output capacity is 131,000 tokens in both input modes, and the reasoning budget is capped at 262,000 tokens. For performance, the model supports rate limits of 2 million tokens per minute and 15,000 requests per minute. The pricing structure for the hosted API includes $2.00 per 1 million input tokens and $6.00 per 1 million output tokens. Implicit cache reads are priced at $0.25 per 1 million tokens. The creation of an explicit cache costs $2.50, while explicit cache reads are $0.17 per 1 million tokens. Notably, cached input is eight times more cost-effective than fresh input, indicating that prefix stability is a more significant factor in cost management than prompt length alone. The Qwen team's release of Qwen3.8-Max underscores the rapid advancements in large language models, particularly in the realm of multimodal capabilities and massive parameter counts.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next