Interestana
Home/News/DeepSeek, Alibaba Launch New AI Models
Bloomberg Markets3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

DeepSeek, Alibaba Launch New AI Models

Chinese artificial intelligence company DeepSeek has launched its latest large language model, DeepSeek-V2, aiming to compete with leading global AI developers. This release, detailed in a company blog post on April 15, 2024, signifies a significant advancement in China's AI capabilities. DeepSeek-V2 reportedly achieves performance comparable to models like Anthropic's Claude 3 Opus and OpenAI's GPT-4 while utilizing substantially fewer computational resources. Specifically, DeepSeek-V2 is claimed to require only 20% of the GPU memory needed by comparable models, a critical factor for efficient deployment and scaling. The model's architecture incorporates a novel Mixture-of-Experts (MoE) approach, which allows for greater efficiency by activating only relevant parts of the neural network for specific tasks. This innovation is key to its reduced memory footprint and enhanced processing speed.

Alibaba Cloud, the cloud computing arm of Chinese e-commerce giant Alibaba, has also introduced its new AI model, Qwen1.5. Announced on April 18, 2024, Qwen1.5 is available in various sizes, with the largest version boasting 110 billion parameters. This model demonstrates strong performance across a range of benchmarks, including coding, mathematics, and general knowledge, positioning it as a formidable competitor in the rapidly evolving AI landscape. Alibaba's move underscores the intense competition within the AI sector, both domestically in China and on the global stage. The development and release of these advanced models by Chinese companies reflect the nation's strategic focus on becoming a leader in artificial intelligence technology. The availability of these powerful models is expected to drive innovation and adoption of AI solutions across various industries within China and potentially beyond.

The competitive landscape for large language models is increasingly crowded, with companies worldwide investing heavily in research and development. Anthropic, known for its Claude series of models, and OpenAI, the creator of the GPT series, have set high benchmarks for performance and capability. DeepSeek-V2's claim of superior efficiency, requiring significantly less GPU memory, addresses a major bottleneck in deploying large AI models. This could make advanced AI more accessible and cost-effective for a wider range of businesses and applications. The MoE architecture, a key feature of DeepSeek-V2, is a growing trend in AI research, enabling models to scale more effectively without a proportional increase in computational cost. The success of this approach in DeepSeek-V2 could influence future model development across the industry.

Alibaba's Qwen1.5, with its substantial parameter count and broad capabilities, further intensifies this competition. The model's performance on benchmarks suggests it can handle complex tasks, making it a valuable tool for developers and enterprises. The release of Qwen1.5 by Alibaba Cloud also highlights the strategic importance of AI for major technology companies seeking to expand their cloud services and offer cutting-edge AI solutions to their customers. As these new models become more widely available and tested, their impact on the global AI market will become clearer, potentially shifting the balance of power among leading AI developers and fostering new waves of AI-driven innovation.

Original source — read the full reporting at the publisher:

Read on Bloomberg Markets

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next