By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Cognition Releases SWE-2 Coding Model, Matching Fable 5.1 at Lower Cost
Cognition, the developer of the Devin AI coding agent, has launched SWE-2, its latest and most advanced coding model. This new model is post-trained using reinforcement learning techniques applied to Kimi K3, an open-source large language model developed by Moonshot AI, which boasts 2.8 trillion parameters. Cognition reports that SWE-2 achieved a score of 50.0% on the FrontierCode 1.1 Main benchmark. This performance is notably close to Fable 5.1, a leading coding model, with only a 1-point difference, while operating at a 64% lower cost. SWE-2 also introduces selectable reasoning-effort levels, a feature developed within a single reinforcement learning run, allowing users to adjust the model's computational intensity and performance trade-offs.
SWE-2 is built upon the foundational infrastructure and methodology established by its predecessor, SWE-1.7, which was similarly post-trained but utilized Kimi K2.7 as its base model. For SWE-2, Cognition scaled its reinforcement learning process to accommodate models with nearly three times the parameters of the previous iteration. The company states that its reinforcement learning approach continues to find significant performance improvements, even on top of the Kimi K3 model, adding an estimated 5 to 6 percentage points to scores across various benchmarks. A key innovation in SWE-2 is its reinforcement learning algorithm, which enables the training of multiple effort levels simultaneously within a single training session. Each effort level is associated with a specific cost penalty, allowing the model to optimize its cost-performance frontier holistically.
In terms of benchmark performance, SWE-2 demonstrates strong results across several evaluations. On FrontierCode 1.1 Main, it scored 50.0%, surpassing Kimi K3's 44.2% and Grok's 48.0%. It also achieved 73.0% on DeepSWE 1.1, outperforming Kimi K3 (68.5%) and Fable 5.1 (67.4%). On Terminal-Bench 2.1, SWE-2 reached an impressive 92.8%, exceeding Kimi K3 (88.3%) and Fable 5.1 (91.4%). Cognition claims that SWE-2 comes within a few percentage points of GPT-6 Astra's performance but at approximately one-quarter of the cost. However, the model shows a comparative weakness on Terminal-Bench 4, where it scored 27.3%, lagging behind Fable 5.1 (55.8%) and GPT-6 Astra (57.9%) by a significant margin. It is important to note that FrontierCode is a proprietary benchmark developed by Cognition, and the reported scores for rival models were also generated by Cognition's internal evaluations.
SWE-2 is currently not available for independent deployment on users' own infrastructure, as it does not offer open weights or a standalone API. Access to SWE-2 is exclusively through Cognition's Devin platform, including Devin: Desktop and CLI versions, with Devin Web and Fusion interfaces slated for future rollout. This integrated approach ensures that the model's capabilities are utilized within Cognition's ecosystem.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.