HeadlinesBriefing favicon HeadlinesBriefing.com

Cognition SWE-2 Hits 92.8 on Terminal-Bench 2.1

Hacker News •
×

Cognition's SWE-2, built on the 2.8T parameter Kimi K3 base with 104B active per token (MoE), achieves 92.8 on Terminal-Bench 2.1. It's Cognition's first multi-trillion-parameter RL scaling.

The serving stack uses MoE inference on NVFP4 and FP8 kernels with quantization-aware training. A draft model retrained with Spec Forge gives 15% longer accept lengths, and a prefill delayer lifts TPM per GPU and tokens/sec per request by 10–20%.

Effort levels: mean steps per run 53 (medium), 80 (high), 98 (max), vs 127 for SWE-1.7. Medium posts a higher Frontier Code score than SWE-1.7 with 58% fewer turns and 81% lower average cost, landing its first real edit at a median of step 18 (SWE-1.7: 48).

Benchmarks: Frontier Code 1.1 Main 50.0, Deep SWE 1.1 73.0, Terminal-Bench 2.1 92.8, Terminal-Bench 4.0 27.3. The 50.0 is one point behind Claude Fable 5.1 (50.9) and 3.3 behind GPT-6 Astra (53.3), at a claimed 64% lower cost than Fable 5.1. Terminal-Bench 4.0 is the soft spot: 27.3 trails Fable 5.1 (55.8) and Astra (57.9).

Proprietary weights, no local run. Available in Devin Desktop and CLI, with rollout on Devin Web and Fusion. No per-token API; cost-per-task comparisons are the pricing surface. All figures are Cognition's own, pending independent replication.