One point. That’s the entire gap between Cognition’s new SWE-2 coding model and Claude Fable 5.1 on FrontierCode 1.1 Main: 50.0% versus 50.9%. The price gap runs the other way at 64%.
If you buy frontier coding capacity by the token, those two numbers should force a rerun of your budget math. A startup took a model anyone can download for free, ran its own reinforcement learning on top, and landed a point behind Anthropic’s flagship while charging about a third of the price. Cognition announced SWE-2 on its blog on September 14, and the pitch is refreshingly blunt: frontier-level agentic coding without frontier pricing.
What shipped, in numbers
The benchmark table is the whole argument, so here it is without decoration.
- FrontierCode 1.1 Main: 50.0% — one point behind Claude Fable 5.1 (50.9%) at 64% lower cost, beating Grok 4.6 on both score and price
- DeepSWE 1.1: 73.0% — level with GPT-5.6 Sol (72.7%) and closing on GPT-6 Astra (74.1%)
- Terminal-Bench 2.1: 92.8% — ahead of everything else in the comparison, GPT-6 Astra included
- Terminal-Bench 4: 27.3% — against Fable 5.1’s 55.8%
Read that last line twice, because it matters as much as the rest. Long-horizon terminal work still belongs to the big labs. If your agents grind through multi-hour shell sessions, SWE-2 is not your answer yet, and paying Fable prices for that capability is defensible. For bounded coding tasks, the price justification just got harder to write.
The base model is the story
Here’s the part I keep thinking about. SWE-2 is post-trained from Kimi K3, the 2.8-trillion-parameter open-weights model from Moonshot. Cognition’s RL recipe added five to six points on top of K3’s own heavily post-trained baseline, on multiple benchmarks.
Two things follow from that. First, open-weight models are now legitimate substrates for commercial frontier products, not a hobbyist tier you benchmark before buying the real thing. Second, there was still measurable headroom in a model everyone assumed had been squeezed dry. Post-training, done well, keeps paying rent.
The technical novelty behind the gains is almost aggressively unglamorous: a single RL run that trains every reasoning-effort level at once, applying a linear cost penalty per effort level tuned to the local slope of the Pareto frontier. Labs typically ship separate mini, standard, and max training runs. Cognition advances the entire cost-performance curve in one pass. Insights like this compound. Every improvement to the recipe lifts every tier you sell.
Meanwhile, DeepSeek couldn’t be killed
The same week delivered the demand-side version of the same lesson. DeepSeek had scheduled V4 Pro for retirement after September 14. Days before the shutdown, the company reversed itself “in response to user demand,” per its API docs, with billing unchanged. The legacy V4-Flash models did get retired, with requests auto-routing to V4.1-Flash at Flash pricing.
API providers almost never walk back a deprecation. That they did says V4 Pro carries real production workloads. The people running those workloads pushed back hard enough to flip a decision days before it took effect.
DeepSeek also slipped out a new surface worth a look: DeepSeek Harness, in developer preview, aimed at agent tool builders. The API now speaks both OpenAI and Anthropic formats, which means Claude Code, GitHub Copilot, and OpenCode can point at DeepSeek as a backend with no code changes. File that under switching costs.
Put the two stories side by side and you get one trend. Pricing power at the frontier is eroding from both ends. Smarter post-training on free open-weight bases (Cognition) presses down from the top. Stubborn open models embedded in production (DeepSeek) hold up the bottom. The middle is where the margin used to live.
What to do with this
Four moves, in order of effort.
- Rerun your make-or-buy math. If a meaningful share of your coding workload is bounded tasks and Terminal-Bench 2.1-class work, a 64% cost delta on near-parity scores is a real line item, not a rounding error. Get the split between bounded and long-horizon work before you do anything else.
- Benchmark on your own repo. Public leaderboards rank models in general. Your codebase decides in particular. The 27.3% versus 55.8% spread on Terminal-Bench 4 shows how violently tiers can separate by task class. A one-day internal eval beats a month of vendor decks.
- Treat open weights as procurement, not ideology. Kimi K3 to SWE-2 is the proof of concept that a free base plus a proprietary recipe ships a commercial product. You may not post-train your own model, but your vendors increasingly do, and their cost structure will show up in their pricing.
- Lower your switching costs on purpose. Dual-format APIs like DeepSeek’s turn backend swaps into config changes. The cheaper it is to move between models, the more of that 64% delta you actually capture when the next SWE-2 arrives. And there will be a next one.
The asset moved
The uncomfortable summary for anyone selling model access: the strategic asset is no longer the base model. Kimi K3 is downloadable. What Cognition owns is the recipe, the RL pipeline, the evals, the data. Those don’t ship in a weights file.
The counterargument is sitting right there in the benchmark table. Terminal-Bench 4, at 27.3%, is where frontier pricing still buys something nobody else can deliver. That’s the honest state of the market in September 2026: the premium survives at the horizon, and everywhere else it’s negotiating with a startup that started from a free model.


