Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra
Points and comments are a snapshot, not live.
SWE-2 achieves 50.0% on FrontierCode 1.1, rivaling top models at 64% lower cost.
Cognition released SWE-2, a coding model post-trained from Kimi K3's 2.8T parameters. It scores 50.0% on FrontierCode 1.1 Main, within one point of Fable 5.1 while being 64% cheaper. Key innovations include Pareto-informed cost penalties in RL, a length-weighted reward baseline, and scaling RL to the multi-trillion-parameter regime. SWE-2 outperforms SWE-1.7 and Grok 4.6 on both score and cost, matches GPT-5.6 Sol and Fable 5/5.1 at lower price, and comes within a few points of GPT-6 Astra at quarter cost. It shows focused exploration, better test coverage, and verification discipline. Available in Devin Desktop, CLI, Web, and Fusion.
What commenters are saying
Commenters are highly skeptical of benchmark gaming. The large gap between Terminal-Bench 2.1 (92.8%) and the newer Terminal-Bench 4 (27.3%) is seen as evidence of overfitting to older benchmarks. Several note the model is distilled from Kimi K3, owned by a Chinese lab, and question Cognition's technical novelty. Users complain about the closed platform (Devin CLI only) and lack of API access. One positive user reports good experience with SWE-1.5 and calls Cognition an underrated player, but is downvoted.