Home Pricing DeepSeek

DeepSeek API Pricing (August 2026)

Last updated:

How much does DeepSeek cost? DeepSeek V4 Flash 0731 costs $0.22 per 1M input tokens and $0.66 per 1M output tokens. DeepSeek V4 Pro 0813 costs $0.66 input and $1.98 output. Cache-hit input on Flash falls to $0.007, which keeps DeepSeek among the cheapest serious APIs at today's posted rates.

DeepSeek Models — Price per 1M Tokens

Relative priceLowestLowerHigherHighestAutomatic log scale across current offers; input, cache and output are graded separately. Color shows cost, not quality.
Showing 2 grouped models from 2 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

DeepSeek V4 pricing: current vs August 16 rates

New rates take effect August 16, 2026 at 4:00 PM UTC. Prices are USD per 1M tokens.

ModelRate periodInputCachedOutputInput / output vs current
DeepSeek V4 Flash 0731Current$0.22$0.0070$0.66Baseline
DeepSeek V4 Flash 0731Off-peak$0.22$0.0070$0.66No change / No change
DeepSeek V4 Flash 0731Peak$0.44$0.014$1.32100% higher / 100% higher
DeepSeek V4 Pro 0813Current$0.66$0.022$1.98Baseline
DeepSeek V4 Pro 0813Off-peak$0.66$0.022$1.98No change / No change
DeepSeek V4 Pro 0813Peak$1.32$0.044$3.96100% higher / 100% higher

Peak windows: 01:00-04:00 and 06:00-10:00 UTC; all other hours are off-peak. Built from the scheduled rate card in the canonical pricing API.

July 31 release update

DeepSeek upgraded the existing deepseek-v4-flash route in place to DeepSeek-V4-Flash-0731. The public-beta snapshot keeps the same architecture, model size, and list prices while adding stronger agent performance, native Responses API support, a 1M-token context window, and up to 384K output. Read the 0731 pricing and benchmark analysis.

August 3 buyer watch

A Hacker News workflow routes coding through DeepSeek V4 Flash, images through GPT-5.6 Luna, and search through Antigravity CLI. It is a community integration rather than a new DeepSeek release, and DeepSeek still describes 0731 as public beta. Read the oh-my-pi stack cost guide.

August 13 pricing update

DeepSeek has published the rate card behind its earlier increase warning. Peak and off-peak billing starts at 16:00 UTC on August 16; off-peak is half the new peak rate but remains above today's active price in every listed V4 billing category. Read the DeepSeek peak-pricing impact analysis.

August 10 local-vs-API buyer watch

An OpenCode builder says the average Go user consumed $1.14 per day of DeepSeek V4 Flash at public rates. Dividing a stated $10,000 dual-DGX setup by that usage produces a 24-year simple payback, but the comparison omits power, operations, utilization, residual value, and privacy benefits. Read the audited DeepSeek API-versus-local break-even.

Prefer a managed multi-model API instead of separate provider accounts? Novita is one option to quote for open-model access. For sustained high-volume workloads, compare it against the GPU break-even math.

Affiliate disclosure: this sponsored link may earn us a commission. It does not affect DeepSeek table order or pricing claims.

DeepSeek is still the budget disruptor in 2026, but the lineup changed. The current official pricing page centers on the V4 family, with Flash as the aggressive low-cost option and Pro as the heavier premium tier.

Current DeepSeek API Pricing

DeepSeek uses cache-hit pricing on repeated prefixes and now splits the family into a cheaper Flash tier and a more expensive Pro tier:

  • DeepSeek V4 Flash 0731 — $0.22/M input, $0.66/M output. Cache hits drop to $0.007/M input.
  • DeepSeek V4 Pro 0813 — $0.66/M input, $1.98/M output, with cached input at $0.022/M.

What changed in V4 Flash 0731?

DeepSeek says the July 31 snapshot is a re-post-training update rather than a new architecture. The stable API model name remains deepseek-v4-flash, so existing integrations receive the new snapshot without changing the route. The release adds native Responses API support and is specifically adapted for Codex-style agent workflows. DeepSeek reports 82.7 on Terminal Bench 2.1, 54.2 on NL2Repo, and 70.3 on Toolathlon Verified, although two other published scores use internal DeepSeek test sets.

Peak and off-peak pricing starts August 16

DeepSeek has now replaced its earlier general increase warning with a complete scheduled rate card. Peak windows are 01:00–04:00 and 06:00–10:00 UTC; all other hours are off-peak. Off-peak rates are 50% below the new peak tier, but the generated schedule table shows they still exceed today's active prices. The current table remains valid until the announced 16:00 UTC cutover.

Labs status

Both current DeepSeek routes are already represented in AI Pricing Guru Labs. The latest accepted run scored V4 Flash and V4 Pro at 48/49 with zero API errors; the Flash result remains labeled pre-0731 because its launch-day refresh failed for insufficient credit and was rejected. This warning changes neither model behavior nor today's active rates, so it does not justify a new inference run. When the new prices activate, Labs costs will be recalculated against the accepted token counts.

The Price Comparison That Matters

Here's how DeepSeek stacks up against the competition:

ModelInput/1MOutput/1Mvs DeepSeek
DeepSeek V4 Flash 0731$0.22$0.66
GPT-5.6 Luna$0.2$1.201x more expensive
Claude Opus 5$5.00$25.0023x more expensive
Gemini 3.1 Pro$2.00$12.009x more expensive

Automatic Caching: The Hidden Advantage

Unlike other providers where you must manually enable caching, DeepSeek's disk-based system caches automatically. Repeated prompts hit the cache without any code changes. Cache hit input costs drop to $0.0028/M tokens — making DeepSeek especially inexpensive for repetitive workloads.

Trade-offs to Consider

DeepSeek isn't perfect for every use case:

  • Rate limits can be tighter than OpenAI/Google during peak hours
  • Content filtering differs from Western providers
  • No free tier like Google Gemini offers
  • Fewer model options — just two models vs OpenAI's 13+

For most developers building cost-sensitive applications, these trade-offs are worth the 90% savings.

DeepSeek on Third-Party Hosts

You can also access DeepSeek models through inference providers like Groq, Together AI, and Fireworks — often with faster inference speeds but at higher prices. Check our full pricing comparison to compare.

Price History

Only models with a recorded price change are charted here.

DeepSeek V4 Flash 0731

DeepSeek V4 Pro 0813

Price history tracking started April 2026. Flat model charts stay hidden until a price change is detected.
View pricing changelog →

Frequently asked questions

How much does DeepSeek cost per token?

DeepSeek V4 Flash 0731 costs $0.22 per 1M input tokens and $0.66 per 1M output tokens. DeepSeek V4 Pro 0813 is $0.66 input and $1.98 output per 1M. Cache-hit input on Flash drops to $0.007 per 1M.

Is DeepSeek raising its API prices?

Yes. DeepSeek will introduce peak and off-peak V4 pricing at 16:00 UTC on August 16, 2026. Off-peak rates are half the new peak rates, but every published off-peak token rate is still higher than the currently active rate.

Does DeepSeek have a free tier?

DeepSeek does not offer a permanent free API tier. New accounts sometimes receive promotional balance, but production usage is pay-as-you-go.

How does DeepSeek compare to OpenAI?

DeepSeek V4 Flash 0731 at $0.22/$0.66 per 1M is about 1x cheaper than GPT-5.6 Luna on input and 2x cheaper on output. The trade-off is a smaller ecosystem and less battle-tested tooling than OpenAI.

What models does DeepSeek offer?

DeepSeek currently exposes two primary production models on its pricing page: DeepSeek V4 Flash 0731 and DeepSeek V4 Pro 0813. Legacy compatibility aliases like deepseek-chat and deepseek-reasoner now map to the V4 family.

What's DeepSeek's context window?

DeepSeek lists a 1M-token context window and up to 384K max output for the current V4 family, which puts it in the long-context tier rather than the old 128K class.

Does DeepSeek have automatic caching?

Yes. DeepSeek applies cache-hit pricing automatically on repeated prefixes. On DeepSeek V4 Flash 0731, cached input drops to $0.007 per 1M tokens.

Methodology

Pricing sourced from DeepSeek's pricing documentation and version details from its official changelog on . All prices in USD per 1 million tokens. Raw data: /api/pricing.json. API docs.

Compare All Providers

See how DeepSeek's pricing compares to every major AI API provider.

Further reading: DeepSeek vs ChatGPT pricing · Cheapest AI APIs in 2026.