DeepSeek R1
DeepSeek R1 is an open-weight reasoning model trained primarily via reinforcement learning. It reproduced much of the performance of OpenAI’s o1 at a fraction of the inference price.
Released
January 20, 2025
Type
reasoning
License
open-weight
Context
128,000 tokens
Pricing
Input
$0.55 / 1M tokens
Output
$2.19 / 1M tokens
Benchmarks
| AIME 2024 | 79.8% |
| GPQA Diamond | 71.5% |
Capabilities
Architecture
Parameters: 685B MoE (37B active)
Links
DeepSeek R1 in the news
Fortune · Jul 25, 2026
DeepSeek said to tell backers of funding pause after viral posts
DeepSeek, a Chinese artificial intelligence company, has informed prospective investors that it is temporarily suspending its second fundraising round. This decision comes days after comments widely attributed to its founder, Liang Wenfeng, regarding US-Chinese AI competition gained significant traction online. The company verbally communicated to some potential backers that investment agreements would not be finalized as anticipated in the coming days. DeepSeek may consider resuming the fundraising process at a later date, according to sources familiar with the matter who requested anonymity to discuss private deliberations. The suspension appears to be partly driven by Liang's dissatisfaction with online reports concerning his remarks made to investors during the company's first financing deal. That initial round, which closed in June, successfully raised $7 billion for the AI lab. Bloomberg has not independently verified the authenticity of the viral posts, which reportedly included a transcript from a meeting Liang conducted with undisclosed individuals. Chinese media outlets, including Yicai, reported this week that the billionaire founder discussed the reliance on Nvidia Corp. chips for AI development and China's ongoing deficit in AI sophistication compared to the United States. DeepSeek has not yet responded to an emailed request for comment regarding the transcript or the fundraising pause. Negotiations are described as fluid, and the company might still decide to proceed with the deal. It remains unclear whether DeepSeek has informed all potential investors in the current round of its intentions. The current funding talks commenced only weeks after DeepSeek concluded its record-breaking first financing round, which attracted prominent investors such as Tencent Holdings Ltd. and Contemporary Amperex Technology Co. Ltd. Prior to the suspension, DeepSeek was reportedly aiming to secure at least 10 billion yuan in additional funds through this follow-on deal, with the final amount potentially increasing based on investor participation. The startup had been targeting a pre-money valuation of at least 480 billion yuan, according to Bloomberg's previous reporting.
Fortune · Jul 22, 2026
As Washington panics about Chinese AI, Jensen Huang says open-source models like Kimi are ‘excellent’ and should be embraced, not banned
Nvidia CEO Jensen Huang urged Washington to embrace open-source AI models developed by Chinese companies, rather than restricting them due to perceived threats. Speaking to Axios on Tuesday, Huang stated that these highly capable, low-cost models, including Moonshot AI's Kimi K3, DeepSeek, and Alibaba's offerings, are "excellent" and should be utilized. This stance comes as the White House reportedly considers restricting U.S. companies' use of Chinese AI models, driven by concerns over potential surveillance and the impact on domestic AI companies. The White House has allegedly considered executive actions to impose conditions on U.S. firms using these models, such as mandating security guarantees and liability acceptance in case of breaches. Huang dismissed the notion that these models could serve as a backdoor for the Chinese government, calling it a "misconception." He emphasized that these models are downloadable and their security guardrails are customizable. Huang also expressed confidence that these open-source models will not surpass American AI advancements, arguing for the necessity of both open-source and closed-source AI systems. He believes that excellent open-source models should be integrated into the AI ecosystem, alongside proprietary models from companies like OpenAI and Anthropic. U.S. companies, including AI coding startup Cursor, are reportedly turning to these Chinese open-source alternatives to mitigate the high costs associated with American AI models. Nvidia did not immediately respond to requests for comment.
Fast Company · Jul 20, 2026
This new, Beijing-based AI model is causing such a stir that subscriptions are now on hold
Moonshot AI, a Beijing-based artificial intelligence company, announced on Sunday that it has temporarily suspended new subscriptions for its Kimi K3 AI model. This decision follows an unprecedented surge in demand that has overwhelmed the company's current capacity within days of the model's launch. The company stated in an X post that "Kimi K3 has received far more love than we expected," and that demand over the past 48 hours has pushed capacity to its limits. Moonshot AI is prioritizing existing subscribers and is working to add capacity, with plans to reopen new subscription spots in batches. The Kimi K3 model, which aims to challenge advanced AI models from companies like Anthropic and OpenAI, has generated significant attention in the U.S. tech industry. It is reportedly the world's largest open-source AI model, boasting 2.8 trillion parameters, a metric used to measure an AI model's capabilities. This rapid ascent and high demand underscore the growing competitiveness of Chinese AI models on the global stage, particularly as the U.S.-China tech race intensifies. Chinese open-source AI models, such as DeepSeek's latest V4, have been gaining international recognition for their advanced capabilities and lower costs. Lian Jye Su, a chief analyst at Omdia, commented that "New model releases generally trigger massive interest, which can strain existing compute infrastructure." He suggested that the demand surge indicates Moonshot AI may not have fully anticipated K3's popularity or possess sufficient compute chips to meet the immediate demand. The Kimi K3 model is described as "very demanding" in terms of its compute requirements, further emphasizing the infrastructure challenges faced by AI developers when dealing with rapid user growth and the need for substantial computational resources.
MarkTechPost · Jul 19, 2026
Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2: Open Trillion-Scale MoE Models Compared on Benchmarks, License, and Serving Cost
Three Chinese research labs have released large, open-weight Mixture-of-Experts (MoE) models that now lead the open leaderboard. Moonshot AI's Kimi K3, DeepSeek V4 Pro, and Zhipu AI's GLM-5.2 all feature million-token context windows and are designed for long-horizon coding and agent tasks. This comparison evaluates them based on measured capability, licensing terms, and the practical cost of serving them. Kimi K3, released on July 16, 2026, is a 2.8-trillion-parameter Stable LatentMoE model. It activates 16 out of 896 experts per token, with Moonshot AI not disclosing the exact active parameter count. Kimi K3 includes native vision and video capabilities, a 1 million-token context window, and always-on reasoning, positioning itself as the first open 3-trillion-parameter class model. DeepSeek V4 Pro, released on April 24, 2026, is a 1.6-trillion-parameter MoE model with 49 billion active parameters, utilizing 384 routed experts plus one shared expert. It also offers a 1 million-token context window with a maximum output of 384,000 tokens. A smaller variant, DeepSeek V4 Flash, has 284 billion total parameters and 13 billion active parameters for more cost-effective workloads. GLM-5.2, released by Zhipu AI on June 13, 2026, is a 744-billion-parameter MoE model with approximately 40 billion active parameters and a 1 million-token context window. It provides High and Max reasoning modes and is available via API access. While GLM-5.2 is the smallest of the three by total parameters, it held the top spot in the open-weight field before Kimi K3's release. The article notes that vendor-reported benchmark scores for these models use different methodologies, making direct comparison challenging. The comparison focuses on three key decision-making axes for AI teams: capability, license, and serving cost. Kimi K3 and DeepSeek V4 Pro are described as 'trillion-parameter' models, with Kimi K3 at 2.8T and DeepSeek V4 Pro at 1.6T. GLM-5.2, with 744 billion total parameters, is the smallest but remains competitive due to its performance and features. The models' licensing terms and the associated infrastructure costs for deployment are critical factors for adoption and practical use.
Fortune · Jul 17, 2026
Businesses are experimenting with cheaper Chinese AI models as U.S. rivals get more expensive
Businesses are increasingly turning to less expensive artificial intelligence models developed by Chinese companies as the cost of leading U.S. AI services rises. This shift is driven by the need to manage escalating expenses associated with token usage and AI operations, prompting some consumer-facing companies to explore open-source alternatives from China. DoorDash, for instance, is experimenting with DoorDash CLI, an AI agent tool that can be accessed via terminal. Andy Fang, DoorDash's co-founder and CTO, stated on X this week that Moonshot AI's models offer superior quality at a lower cost. Moonshot AI is not the only Chinese AI provider gaining traction. Cursor, an AI coding startup, utilized Moonshot's Kimi model to develop its Composer 2 coding agent. Additionally, the startup Lindy has reportedly transitioned from using tools provided by Anthropic to DeepSeek's V4 models, according to the Financial Times. These companies are joining larger entities such as Airbnb and Siemens, which are also investigating the integration of Chinese AI providers like Alibaba and DeepSeek to mitigate rising AI expenditures. According to Yasir Atalan, deputy director and data fellow at the Center for Strategic and International Studies, this trend is influenced by three key factors: cost, capability, and the availability of open-source models. Atalan told Fortune that high-performance models from U.S. companies appear costly in comparison to their Chinese counterparts. The appeal of open-source models is particularly strong for countries outside the U.S., as it allows them to avoid sharing sensitive enterprise data.