DeepSeek V4 Flash 0731 Pricing and Benchmark Analysis
DeepSeek V4 Flash 0731 improved agent results at launch. See official rates, independent benchmarks, and its warning of a future price increase.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- DeepSeek upgraded the stable deepseek-v4-flash API route in place to the 0731 snapshot; existing API calls do not need a new model name.
- Launch-day list prices did not change, but DeepSeek now warns that a significant overall API increase is coming without giving rates or timing.
- Artificial Analysis scores 0731 at 49.93 versus 40.28 for the earlier Flash snapshot, but its evaluation used 210 million output tokens.
- Our fresh Labs rerun is blocked by exhausted OpenRouter credits, so the prior 48/49 result is explicitly labeled pre-0731.
DeepSeek V4 Flash versus nearby intelligence leaders
USD per 1M tokens. Input and output rates are charted separately.
Estimate the cost of a production canary
Assumes 75% input tokens and 25% output tokens using current per-million rates.
DeepSeek V4 Flash 0731
deepseek
$3.30
- Input share
- $1.65
- Output share
- $1.65
GPT-5.6 Luna
openai
$4.50
- Input share
- $1.50
- Output share
- $3.00
DeepSeek V4 Pro 0813
deepseek
$9.90
- Input share
- $4.95
- Output share
- $4.95
Gemini 3.6 Flash
$15.00
- Input share
- $5.63
- Output share
- $9.38
DeepSeek V4 Flash 0731 and benchmark-neighbor API prices
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| DeepSeek V4 Flash 0731 | deepseek | $0.22 | $0.0070 | $0.66 |
| DeepSeek V4 Pro 0813 | deepseek | $0.66 | $0.022 | $1.98 |
| GPT-5.6 Luna | openai | $0.2 | $0.02 | $1.20 |
| Gemini 3.6 Flash | $0.75 | $0.075 | $3.75 |
Built from pricing.json at publish time.
DeepSeek has moved DeepSeek-V4-Flash-0731 into public beta without creating a new API route or raising its regular token prices. The existing deepseek-v4-flash name now serves the re-post-trained checkpoint. DeepSeek also published the model weights under the MIT license. Its architecture and size are unchanged; V4 Pro and DeepSeek’s app and web models were not upgraded.
For buyers, the launch matters because it improves the low end of the coding-agent cost curve. See the live DeepSeek pricing page, full API price comparison, and token cost calculator for current values.
DeepSeek V4 Flash 0731 pricing
DeepSeek lists a 1 million-token context window, maximum output of 384,000 tokens, thinking and non-thinking modes, and concurrency of 2,500. It supports Chat Completions, the Anthropic-compatible interface, tools, JSON, and the Responses API.
The live table and chart above compare Flash with V4 Pro and nearby intelligence leaders using today’s canonical data. DeepSeek did not change Flash’s regular token rates on July 31, but token price alone does not prove the same cost per completed task.
DeepSeek’s pricing page now warns that it plans a significant overall API price increase, but it does not disclose the new rates, affected billing items, or effective date. AI Pricing Guru has therefore kept the canonical dataset on the current published rate card. See our rapid-response price increase analysis for the confirmed details and buyer checklist.
What DeepSeek reports improved
DeepSeek’s changelog calls out agent and coding results rather than a broad general-knowledge scorecard:
| Benchmark | DeepSeek-reported 0731 score | Caveat |
|---|---|---|
| Terminal Bench 2.1 | 82.7 | Public benchmark; DeepSeek harness minimal mode, max effort |
| NL2Repo | 54.2 | Vendor-reported run |
| Cybergym | 76.7 | Vendor-reported run |
| DeepSWE | 54.4 | Vendor-reported run |
| Toolathlon Verified | 70.3 | Vendor-reported run |
| Agent Last Exam | 25.2 | Vendor-reported run |
| Automation Bench Public | 25.1 | Vendor-reported run |
| DSBench FullStack | 68.7 | Internal DeepSeek test set |
| DSBench Hard | 59.6 | Internal DeepSeek test set |
DeepSeek says the code-agent benchmarks used its unreleased minimal harness at maximum effort, top_p=0.95, and temperature 1.0. A tuned private harness may not transfer to another coding agent, lower effort setting, or shorter output budget.
DeepSeek says 0731 is adapted for Codex-style Responses API integrations. The stable model name avoids migration work but also upgrades behavior in place, so regression-test prompts calibrated for the preview.
The open weights create a second deployment path, but DeepSeek’s official vLLM recipe requires a four-GPU GB300 node. Most application teams should compare API economics before self-hosting.
Independent price-performance result
Artificial Analysis independently tested the reasoning, maximum-effort configuration. Its Intelligence Index v4.1 gives DeepSeek V4 Flash 0731 a score of 49.93, up from 40.28 for the earlier V4 Flash snapshot. That is a gain of 9.65 points, or about 24% relative.
The score is nearly level with Gemini 3.6 Flash at 50.07 and higher-effort GPT-5.6 Luna variants. Artificial Analysis ranks 0731 first for weighted cost per Intelligence Index task on its comparison page.
The catch is token use. Artificial Analysis reports 210 million output tokens for the Intelligence Index evaluation, compared with a 62 million median across models on the page. The model is cheap per token, but a long reasoning trace can consume part of that advantage. Procurement should track cost per accepted task, latency, and output volume instead of assuming the lowest rate card always creates the lowest bill.
Artificial Analysis also reports an Omniscience Index of -15.72. That metric rewards correct knowledge and penalizes hallucinated answers without penalizing refusals. It is a warning against reading the coding-agent gains as universal factual reliability.
Labs status: included, but the 0731 refresh is blocked
DeepSeek V4 Flash is already part of the AI Pricing Guru Cost-per-Task Labs leaderboard. The last accepted run scored the previous snapshot at 48 correct answers out of 49 with zero API errors. Because it predates July 31, the public label now says pre-0731.
We attempted a fresh 49-task run at 11:38 UTC on July 31. All requests returned 402 Insufficient credits, so the publication guard rejected the incomplete run.
A valid refresh requires funded credit and confirmation that deepseek/deepseek-v4-flash serves 0731. Until then, labeling the old 48/49 result as 0731 would be misleading.
Buyer verdict
DeepSeek V4 Flash 0731 is a high-priority test for coding agents, repository work, terminal automation, and tool-heavy workflows where cheap tokens allow more verification and retries. The in-place route upgrade makes adoption easy, and the unchanged rate keeps the economic case intact. Compare current alternatives on the Google AI pricing page, OpenAI pricing page, and our DeepSeek vs ChatGPT guide.
It is not an automatic replacement for premium models. The independent run shows high verbosity, the official benchmark configuration depends on a DeepSeek harness that is not yet public, and two headline scores use internal test sets. Teams should compare accepted changes, review findings, rollbacks, latency, and total token volume on their own repositories.
For third-party hosting, compare Novita’s listed DeepSeek routes, but confirm the served snapshot before treating any generic V4 Flash listing as 0731.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored Novita link at no extra cost to you. It does not affect the pricing or benchmark analysis.
The practical move is to canary 0731 on a fixed set of real tasks, cap its reasoning budget, and measure cost per accepted result. If its agent gains hold under your harness, the unchanged rate can buy materially more useful work than the preview. If verbosity or factual reliability dominates, route only the well-scoped coding and tool tasks where Flash earns its cost advantage. For the same-day counterargument to generic prompt routing, read why Manifest deprecated its LLM router.
Sources: DeepSeek’s official July 31 changelog, models and pricing documentation, Codex integration guide, official Hugging Face model card, Artificial Analysis’ corrected DeepSeek V4 Flash 0731 page, and the Hacker News discussion.