Cognition SWE-2 Launch: Pricing Impact
Cognition says SWE-2 nears frontier coding performance at lower task cost. See the benchmarks, availability, pricing caveat, and buyer advice.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- Cognition launched SWE-2 on September 10, 2026 in Devin Desktop and CLI, with Web and Fusion rollout underway.
- The company reports near-frontier coding scores with lower average benchmark task cost, but those percentages are not a public per-token SWE-2 API rate.
- There is no standalone SWE-2 API price or model endpoint in the launch announcement; buyers consume it through Devin products and should compare completed-task cost.
- Use a controlled repository trial before switching: measure accepted fixes, review time, retries, wall-clock time, and total Devin consumption.
Current API cost baseline for SWE-2's benchmark peers
USD per 1M tokens. Input and output rates are charted separately.
Estimate a direct-model coding workload
Assumes 75% input tokens and 25% output tokens using current per-million rates.
Kimi K3
moonshot
$60.00
- Input share
- $22.50
- Output share
- $37.50
GPT-5.6 Sol
openai
$80.00
- Input share
- $30.00
- Output share
- $50.00
Claude Fable 5.1
anthropic
$200.00
- Input share
- $75.00
- Output share
- $125.00
GPT-6 Astra
openai
$200.00
- Input share
- $75.00
- Output share
- $125.00
Public API pricing for models named in Cognition's comparison
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| Claude Fable 5.1 | anthropic | $10.00 | $0.25 | $50.00 |
| GPT-5.6 Sol | openai | $4.00 | $0.4 | $20.00 |
| GPT-6 Astra | openai | $10.00 | $1.00 | $50.00 |
| Kimi K3 | moonshot | $3.00 | $0.3 | $15.00 |
Built from pricing.json at publish time.
Cognition launched SWE-2 on September 10, 2026, calling it the company’s most advanced coding model. It is available in Devin Desktop and CLI, with rollout underway for Devin Web and Fusion.
The headline is cost efficiency, not a new public token rate. Cognition says SWE-2 reached 50.0% on FrontierCode 1.1 Main—within one percentage point of Claude Fable 5.1—while costing 64% less in that benchmark. Those are vendor-run task-cost results, not a standalone SWE-2 API price card.
What launched
SWE-2 is a multi-trillion-parameter coding model post-trained from Kimi K3. Cognition says one reinforcement-learning run trains medium, high, and max reasoning-effort levels using different cost penalties, so users can trade more computation for harder tasks without changing models.
Availability is currently product-led:
| Surface | September 11 status | Billing fact disclosed at launch |
|---|---|---|
| Devin Desktop | Available | Uses Devin consumption; no SWE-2 token rate announced |
| Devin CLI | Available | Uses Devin consumption; no SWE-2 token rate announced |
| Devin Web | Rolling out | No standalone SWE-2 SKU announced |
| Fusion | Rolling out | No standalone SWE-2 SKU announced |
| Public model API | Not announced | No endpoint or input/output rate published |
What the cost claims actually show
Cognition reports these results from three runs over the 100-task FrontierCode 1.1 Main set and companion benchmarks:
| Model | FrontierCode 1.1 Main | DeepSWE 1.1 | Terminal-Bench 2.1 |
|---|---|---|---|
| SWE-2 | 50.0% | 73.0% | 92.8% |
| Kimi K3 | 44.2% | 68.5% | 88.3% |
| Grok 4.6 | 48.0% | 67.5% | 88.4% |
| Claude Fable 5.1 | 50.9% | 67.4% | 91.4% |
| GPT-5.6 Sol | 47.5% | 72.7% | 88.8% |
| GPT-6 Astra | 53.3% | 74.1% | 89.9% |
At medium effort, SWE-2 reportedly used 58% fewer turns and cost 81% less on average than SWE-1.7 while scoring higher on FrontierCode. Cognition also says SWE-2 matches GPT-5.6 Sol and Fable 5/5.1 at a fraction of their task cost and comes within a few points of GPT-6 Astra at about one-quarter of the cost.
These comparisons are useful directional evidence. They are not invoices, independent replications, or equivalent to the live public API rates above. Agent cost depends on tool calls, context reuse, retries, repository size, effort level, and whether a patch survives review.
Who benefits—and who should wait
Teams already using Devin have the clearest reason to test SWE-2. A model that reaches a correct edit sooner can reduce both consumption and developer supervision even when its nominal compute per step is not the lowest.
Wait if you need a direct model API, portable token billing, self-hosted weights, or independently reproduced performance. Cognition did not announce those options. Kimi K3 remains the disclosed base model, but SWE-2’s post-training and agent behavior are the product being sold—not merely access to the base weights.
For direct API alternatives, compare current Anthropic pricing, OpenAI pricing, and xAI pricing. The token calculator estimates model charges, while our best AI for coding guide covers broader workflow tradeoffs.
What developers should do now
- Run the same bounded issue set through SWE-2 medium and your current coding agent.
- Record accepted fixes, failed tests, retries, human review minutes, wall-clock time, and total consumption.
- Escalate only the hardest tasks to high or max effort; Cognition says the behavioral and cost gap between effort levels is material.
- Keep permissions narrow and inspect tool activity. Better benchmark performance does not remove repository, secret, or deployment risk.
- Decide on cost per accepted change—not benchmark score, token rate, or turns in isolation.
If a prototype needs a disposable test environment, compare DigitalOcean development infrastructure. Keep production secrets and customer data outside agent sandboxes until controls are validated.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored DigitalOcean link at no extra cost to you. It does not affect this analysis.
Labs coverage decision
SWE-2 does not enter the AI Pricing Guru Labs leaderboard yet. Cognition has not announced a public inference endpoint or portable request-level billing, so our deterministic text suite cannot run the model under comparable conditions.
A valid coding-agent evaluation needs a fixed repository snapshot, identical issues, clean environments, recorded tool traces, accepted-test scoring, review-time measurement, and complete consumption logs. We will add SWE-2 when reproducible access supports that protocol.
Bottom line
SWE-2 is a serious cost-performance claim for coding agents, but not a new commodity API rate. The launch matters most to Devin users who can test whether fewer turns and stronger patches lower cost per accepted change.
Treat Cognition’s benchmarks as a shortlist signal. Run a controlled trial, preserve the full cost ledger, and wait for a public endpoint before comparing SWE-2 as if it were a normal per-token API model.
Source: Cognition’s official SWE-2 launch announcement, published September 10, 2026. Availability, benchmark figures, and pricing status checked September 11, 2026.