Qwen3 235B
Qwen3 235B is Alibaba’s flagship open-weight model. The MoE architecture and Apache 2.0 license make it a popular base for fine-tuning in the open-source community.
Released
April 29, 2025
Type
llm
License
open-weight
Context
128,000 tokens
Capabilities
Architecture
Parameters: 235B MoE (22B active)
Links
Qwen3 235B in the news
The Hacker News · Sep 11, 2026
Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks
Anthropic announced on Thursday that it has identified and disrupted industrial-scale illicit distillation attacks targeting its Claude AI model. These sophisticated attacks were traced back to seven distinct artificial intelligence laboratories based in China. The identified entities include prominent organizations such as Alibaba, Moonshot, DeepSeek, Z.ai (also known as Zhipu), and MiniMax. The attacks involved knowledge distillation, a legitimate machine learning technique where a larger, more capable AI model acts as a 'teacher' to train a smaller model. However, in this context, the labs were allegedly using Anthropic's powerful Claude models to illicitly train their own AI systems without authorization, bypassing Anthropic's terms of service and potentially infringing on intellectual property rights. Anthropic's security team detected these activities by monitoring for unusual patterns of access and data exfiltration that were indicative of large-scale, systematic attempts to replicate Claude's capabilities. The company stated that these distillation efforts were "industrial-scale," suggesting a significant investment of resources and computational power by the implicated labs. Such attacks aim to extract proprietary knowledge and model parameters from a teacher model to create a student model that mimics its performance, often at a fraction of the training cost and time. By engaging in this unauthorized distillation, the Chinese labs sought to gain a competitive advantage by rapidly developing advanced AI models based on Anthropic's research and development. In response to the discovery, Anthropic took immediate action to mitigate the threat. The company stated that it has taken steps to block the identified actors and prevent further unauthorized access and misuse of its Claude models. This disruption aims to protect Anthropic's intellectual property and maintain the integrity of its AI systems. The company emphasized that knowledge distillation is a valid technique when conducted ethically and with proper licensing, but the actions of these seven labs constituted a violation of their acceptable use policies. The incident highlights the ongoing challenges in securing advanced AI models against sophisticated intellectual property theft and unauthorized replication attempts, particularly in a competitive global AI landscape. Anthropic has not disclosed the specific versions of Claude that were targeted in these attacks, nor the exact methods used by the labs to extract the necessary data for distillation. However, the scale and systematic nature of the attacks suggest a coordinated effort. The company's proactive stance in identifying and reporting these activities underscores the increasing importance of AI security and the need for robust defenses against emerging threats. The involvement of multiple well-known Chinese AI companies also points to the intense competition and rapid development occurring within China's AI sector, where the pursuit of cutting-edge capabilities may sometimes lead to questionable practices. Anthropic's disclosure serves as a warning to other AI developers and a signal of the evolving security landscape in artificial intelligence.
TechCrunch · Sep 10, 2026
Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek
Anthropic detailed allegations of persistent "distillation attacks" originating from Chinese AI companies, including Alibaba, Moonshot AI, and DeepSeek, in a report released on Thursday. These attacks, which involve extracting proprietary information from AI models, have reportedly escalated in recent months amidst intensifying competition within the artificial intelligence sector. Distillation attacks, a form of intellectual property theft, occur when a malicious actor trains a smaller, unauthorized model to mimic the behavior and output of a larger, proprietary model. This process often involves repeatedly querying the target model and using its responses to train the attacker's model, effectively stealing its learned capabilities without direct access to its underlying architecture or training data. The report specifically names Alibaba, a multinational technology conglomerate, Moonshot AI (also known as Kimi AI), a Chinese AI company focused on large language models, and DeepSeek, an AI research organization, as entities allegedly involved in these activities. Anthropic, a leading AI safety and research company known for its Claude series of large language models, asserts that these attacks pose a significant threat to the integrity and competitive landscape of AI development. The company's research indicates a pattern of behavior consistent with distillation, where models developed by these entities exhibit performance characteristics remarkably similar to Anthropic's own proprietary models, often within a short timeframe after the proprietary models' release or significant updates. Anthropic's findings suggest that the sophistication and frequency of these attacks have increased, reflecting a broader trend of heightened competition and a race to develop advanced AI capabilities. The company emphasizes that such practices not only undermine fair competition but also raise concerns about the security and originality of AI models entering the market. By replicating the performance of established models, attackers can potentially bypass the extensive research, development, and computational resources required to build such systems from scratch. This practice can lead to a market flooded with imitative products that may lack the robust safety features, ethical considerations, and genuine innovation of the original models. The implications of these alleged distillation campaigns extend beyond intellectual property concerns. They could impact the trust and transparency in the AI ecosystem, making it difficult for users and businesses to discern the true origins and capabilities of different AI models. Anthropic's report serves as a call for greater vigilance and potentially new industry standards or regulatory measures to address these evolving threats in the rapidly advancing field of artificial intelligence. The company has not yet detailed specific technical evidence in the public report but indicated that further information may be shared through appropriate channels, underscoring the seriousness of the allegations and the potential ramifications for the global AI industry.
Ars Technica · Sep 9, 2026
Six Chinese AI firms accused of aggressively copying US frontier models
Six Chinese artificial intelligence companies have been accused by United States government agencies of engaging in industrial-scale attacks to distill the capabilities of US frontier AI models. The allegations were detailed in a joint release on Tuesday by the National Security Agency (NSA), the Cybersecurity and Infrastructure Security Agency (CISA), and the Federal Bureau of Investigation (FBI). The named companies include DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI. These entities are alleged to have been actively attacking US models since at least late 2024, a practice that could significantly reduce their own development timelines and financial expenditures for training frontier models. The agencies stated that these Chinese firms "likely" acted with "Chinese government awareness" when extracting capabilities from prominent US models. The specific US models targeted include variants of OpenAI's GPT series, Anthropic's Claude, Google's Gemini, and xAI's Grok. The process of "distillation" involves using a powerful "teacher" model to train a smaller, more efficient "student" model, often by having the student model mimic the teacher's outputs. When conducted on a large scale and without authorization, this can be a form of intellectual property theft and a shortcut to developing advanced AI capabilities. The accusations highlight a growing concern within the US regarding the competitive landscape of AI development and the potential for foreign adversaries to gain an unfair advantage. By allegedly bypassing the extensive research, development, and computational resources required to build frontier AI models from scratch, these Chinese firms could accelerate their own progress and potentially reduce the cost of developing advanced AI systems. This practice, if proven, represents a significant challenge to the intellectual property and technological leadership of US-based AI developers. The joint statement from the NSA, CISA, and FBI underscores the national security implications of such alleged activities. The ability of foreign entities to rapidly acquire advanced AI capabilities through illicit means could have far-reaching consequences for economic competitiveness and national security. The agencies' involvement indicates that these concerns are being treated with the utmost seriousness, involving intelligence and law enforcement branches of the US government. The investigation into these claims is ongoing, and further actions may be taken based on the findings.
MarkTechPost · Aug 28, 2026
GLM-5.3-Flash vs Qwen3.8-Flash-Next: Two Chinese AI Labs Independently Converge on the Same Model Architecture
Two leading Chinese artificial intelligence labs, Z.ai and Alibaba's Qwen team, independently developed and released frontier open-weight models with strikingly similar architectures within a single week. Z.ai launched GLM-5.3-Flash, a 320 billion parameter multimodal Mixture-of-Experts (MoE) model featuring 18 billion active parameters. Concurrently, Alibaba's Qwen team unveiled Qwen3.8-Flash-Next, a 125 billion parameter model with 6 billion active parameters, which serves as a preview of the upcoming Qwen4 architecture. Despite their independent development, the configurations of these two models are nearly identical, indicating a potential convergence in cutting-edge AI model design. Both models employ a 3:1 hybrid attention mechanism, combining linear and full attention. They also utilize a compressed indexer to manage context, capped at 2048 tokens, and widen the residual stream into four gated branches. Furthermore, both models were trained using the Muon optimizer, with fused parameter matrices split prior to orthogonalization. GLM-5.3-Flash is the first natively multimodal model in the GLM-5 series, released under the MIT license on Hugging Face. Z.ai tested it anonymously as Ox Alpha on OpenRouter, where it quickly became the most popular model of the week. This model was trained on a substantial 30 trillion token multimodal corpus and supports a context window of 1 million tokens. Z.ai claims that GLM-5.3-Flash outperforms its predecessor, GLM-5.2, across various benchmarks at one-tenth the cost, while achieving performance comparable to Claude Opus 4.8 on coding and agentic tasks. Its listed pricing is $0.15 per million input tokens and $0.50 per million output tokens. Qwen3.8-Flash-Next fulfills a similar role to Qwen3-Next in the Qwen3.5 release, acting as an early public preview of the next generation of Qwen architecture. The model card specifies a main model of 125 billion parameters, supplemented by an additional 51 billion n-gram embedding table, with 6 billion parameters activated per token. Its native context length is 262,144 tokens, which can be extended to 1 million tokens using YaRN. The Qwen team reported that the training of Qwen3.8-Flash-Next required approximately one-ninth of the compute resources needed for Qwen3.7-Plus. The technical report accompanying this release is titled “On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability.” The shared architectural choices suggest a consensus among leading AI research labs on effective strategies for building powerful and efficient large language models, particularly in the realm of multimodal capabilities and context handling. The only notable point of disagreement among the described configurations is a single architectural detail, and one lab has reportedly dissented from the prevailing design choices.
Search Engine Land · Aug 27, 2026
Your images have a new job in AI search
For over a century, images have served as a crucial tool for consumers to make purchasing decisions without direct product interaction. As early as 1897, Sears utilized catalog illustrations to enable shoppers to "order intelligently… as well as if you were in our store selecting the goods from stock," effectively substituting the visual representation for a physical visit. This strategy scaled significantly, with Sears distributing 50 million catalogs annually by 1916. Decades later, the Delia's catalog continued this trend, reaching a peak of 55 million mailings per year. Photo producer Jim Trzaska observed that teenage girls used these catalogs as sales tools, presenting pictures to persuade their parents into clothing purchases. The advent of social media further amplified this visual sales tactic, transforming personal feeds into platforms for product discovery. Currently, the role of images is expanding to include communication with artificial intelligence search engines. An AI answer engine can now analyze a photograph, deconstruct it into mathematical tokens, and interpret its content. This means an image must not only convey meaning to a human viewer but also be comprehensible to a machine. The AI needs to understand what is depicted in the frame, its significance, and whether the surrounding textual content on a webpage corroborates the image's message. This dual requirement fundamentally alters the practice of image search engine optimization (SEO), necessitating that images are optimized for machine readability alongside human appeal. The landscape of search is rapidly evolving towards visual interaction. Google Lens processes approximately 20 billion visual searches monthly, demonstrating the growing reliance on visual input. On Pinterest, visual search constitutes 30% of all searches and exhibits a 62% higher conversion rate compared to text-based searches, as a camera can identify specific items more efficiently than descriptive words. Pinterest Lens alone facilitates 1.5 billion visual searches each month. In China, Alibaba's camera-first shopping platform has consistently exceeded 10 million visual searches daily for years, underscoring the mainstream adoption and monetization of camera-to-cart functionalities. This integration of visual search into e-commerce and information retrieval signifies a paradigm shift, where the image acts as a gateway for both user queries and AI-generated answers. Further illustrating this trend, Google filed a patent in 2023, published in April 2026, which outlines an answer delivery system. In this system, the primary source for an answer is selected based on an image match, with subsequent text from the surrounding page being incorporated to construct the final response. This patent highlights Google's strategic direction towards prioritizing visual information in its search algorithms, reinforcing the idea that images are becoming increasingly integral to how users find information and products online.