Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra

cognition.com

338 points by seelos 8 hours ago


postalcoder - 7 hours ago

If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).

Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"