Interestana
Home/News/Two Chinese AI Labs Independently Design Identical Model Architectures
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Two Chinese AI Labs Independently Design Identical Model Architectures

Two leading Chinese artificial intelligence labs, Z.ai and Alibaba's Qwen team, independently developed and released frontier open-weight models with strikingly similar architectures within a single week. Z.ai launched GLM-5.3-Flash, a 320 billion parameter multimodal Mixture-of-Experts (MoE) model featuring 18 billion active parameters. Concurrently, Alibaba's Qwen team unveiled Qwen3.8-Flash-Next, a 125 billion parameter model with 6 billion active parameters, which serves as a preview of the upcoming Qwen4 architecture. Despite their independent development, the configurations of these two models are nearly identical, indicating a potential convergence in cutting-edge AI model design. Both models employ a 3:1 hybrid attention mechanism, combining linear and full attention. They also utilize a compressed indexer to manage context, capped at 2048 tokens, and widen the residual stream into four gated branches. Furthermore, both models were trained using the Muon optimizer, with fused parameter matrices split prior to orthogonalization. GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, released under the MIT license on Hugging Face. Z.ai tested it anonymously as Ox Alpha on OpenRouter, where it quickly became the most popular model of the week. This model was trained on a substantial 30 trillion token multimodal corpus and supports a context window of 1 million tokens. Z.ai claims that GLM-5.3-Flash outperforms its predecessor, GLM-5.2, across various benchmarks at one-tenth the cost, while achieving performance comparable to Claude Opus 4.8 on coding and agentic tasks. Its listed pricing is $0.15 per million input tokens and $0.50 per million output tokens. Qwen3.8-Flash-Next fulfills a similar role to Qwen3-Next in the Qwen3.5 release, acting as an early public preview of the next generation of Qwen architecture. The model card specifies a main model of 125 billion parameters, supplemented by an additional 51 billion n-gram embedding table, with 6 billion parameters activated per token. Its native context length is 262,144 tokens, which can be extended to 1 million tokens using YaRN. The Qwen team reported that the training of Qwen3.8-Flash-Next required approximately one-ninth of the compute resources needed for Qwen3.7-Plus. The technical report accompanying this release is titled “On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability.” The shared architectural choices suggest a consensus among leading AI research labs on effective strategies for building powerful and efficient large language models, particularly in the realm of multimodal capabilities and context handling. The only notable point of disagreement among the described configurations is a single architectural detail, and one lab has reportedly dissented from the prevailing design choices.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next