Cognition introduced SWE-2, calling it their closest model yet to the frontier: on leading evals it matches recent frontier models at up to ~70% lower cost. They say they scaled RL to multiple trillions of parameters with a refined recipe that pushes the Pareto curve on both capability and cost. SWE-2 already powers products such as Devin Voice.
Key Takeaways
- ✓SWE-2 claims parity with recent frontier models on leading coding evals.
- ✓Up to ~70% lower cost at similar capability; Pareto-focused release.
- ✓RL scaled to multi-trillion parameters; already used in Devin products.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.