Home › Pricing › DeepSeek

DeepSeek API Pricing (September 2026)

Last updated:

How much does DeepSeek cost? DeepSeek V4.1 Flash costs $0.15 per 1M cache-miss input tokens, $0.003 per 1M cache-hit tokens, and $0.6 per 1M output tokens off-peak. Peak rates are exactly double.

DeepSeek Models — Price per 1M Tokens

Relative priceLowestLowerHigherHighestAutomatic log scale across current offers; input, cache and output are graded separately. Color shows cost, not quality.
Showing 2 grouped models from 2 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

DeepSeek current peak and off-peak rates

Rates effective since September 10, 2026 at 4:00 AM UTC. Prices are USD per 1M tokens.

ModelRate periodInputCachedOutputInput / output vs current
DeepSeek V4.1 FlashOff-peak baseline$0.15$0.0030$0.6Comparison baseline
DeepSeek V4.1 FlashOff-peak$0.15$0.0030$0.6No change / No change
DeepSeek V4.1 FlashPeak$0.3$0.0060$1.20100% higher / 100% higher
DeepSeek V4 Pro 0813Off-peak baseline$0.66$0.022$1.98Comparison baseline
DeepSeek V4 Pro 0813Off-peak$0.66$0.022$1.98No change / No change
DeepSeek V4 Pro 0813Peak$1.32$0.044$3.96100% higher / 100% higher

Peak windows: 01:00-04:00 and 06:00-10:00 UTC; all other hours are off-peak. Built from the scheduled rate card in the canonical pricing API.

September 10 release

DeepSeek V4.1 Flash is live on deepseek-flash with native vision and lower prices. The older Flash and Flash Vision experimental names temporarily route to it. DeepSeek initially said V4 Pro would redirect on September 14, but the latest official pricing-page footnote now says Pro will continue with unchanged billing. Read the V4.1 Flash launch and migration analysis.

September 12 local-inference buyer watch

An independent SSD-streaming experiment ran the original V4.1 Flash FP4/FP8 checkpoint on a 16 GB M1 Mac mini, but the optimized result was 108 seconds to first token and 22.8 seconds per token. The API price did not change, and V4.1 Flash remains included in Labs through its managed route. Read the local-versus-hosted buyer analysis.

September 16 cyber-agent benchmark watch

Enclave reports that V4.1 Flash achieved code execution on 11 of 11 vulnerable targets while four fixed controls remained secure. Its audit found six intended exploit paths and five alternate routes in the private test environment; the accepted runs cost $4.65 with more than 99% of input tokens served from cache. The API price did not change. Read the cost, cache, and benchmark-limit analysis.

July 31 release update

DeepSeek upgraded the existing deepseek-v4-flash route in place to DeepSeek-V4-Flash-0731. The public-beta snapshot keeps the same architecture, model size, and list prices while adding stronger agent performance, native Responses API support, a 1M-token context window, and up to 384K output. Read the 0731 pricing and benchmark analysis.

August 3 buyer watch

A Hacker News workflow routes coding through DeepSeek V4 Flash, images through GPT-5.6 Luna, and search through Antigravity CLI. It is a community integration rather than a new DeepSeek release, and DeepSeek still describes 0731 as public beta. Read the oh-my-pi stack cost guide.

August 13 pricing update

DeepSeek introduced peak and off-peak billing on August 16. The time windows remain active, but V4.1 Flash received a lower rate card on September 10. Read the original DeepSeek peak-pricing impact analysis.

August 10 local-vs-API buyer watch

An OpenCode builder says the average Go user consumed $1.14 per day of DeepSeek V4 Flash at public rates. Dividing a stated $10,000 dual-DGX setup by that usage produces a 24-year simple payback, but the comparison omits power, operations, utilization, residual value, and privacy benefits. Read the audited DeepSeek API-versus-local break-even.

Prefer a managed multi-model API instead of separate provider accounts? Novita is one option to quote for open-model access. For sustained high-volume workloads, compare it against the GPU break-even math.

Affiliate disclosure: this sponsored link may earn us a commission. It does not affect DeepSeek table order or pricing claims.

DeepSeek remains the budget disruptor in 2026, and V4.1 Flash now leads its public API story: a new multimodal architecture, a new deepseek-flash route, and a lower time-of-day rate card.

Current DeepSeek API Pricing

DeepSeek uses cache-hit pricing on repeated prefixes and now splits the family into a cheaper Flash tier and a more expensive Pro tier:

  • DeepSeek V4.1 Flash — $0.15/M cache-miss input, $0.6/M output, and $0.003/M cache-hit input off-peak.
  • DeepSeek V4 Pro 0813 — $0.66/M cache-miss input, $1.98/M output, and $0.022/M cached input off-peak.

What changed in V4.1 Flash?

DeepSeek describes V4.1 Flash as a 552-billion-parameter mixture-of-experts model using a new Causal Encoder–Decoder architecture. It activates 8B parameters for input and 16B for output, supports native image input, keeps a 1M-token context window, and allows up to 384K output. These are provider specifications, not independent performance guarantees.

Peak and off-peak pricing

Peak windows are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. All other hours, including weekends, are off-peak. The canonical comparison rate uses the off-peak tier; the generated schedule table exposes both payable tiers so budgets do not quietly understate weekday peak usage.

V4 Pro is not being retired on September 14

The launch announcement and September 11 customer email said deepseek-v4-pro would redirect to V4.1 Flash at 04:00 UTC on September 14. DeepSeek's current pricing page now contains a newer footnote saying it will continue V4 Pro after that date with unchanged billing and provide further notice before any change. We therefore keep Pro active and will monitor the official page daily.

Labs status

V4.1 Flash is a new served model, not a billing-only rename, so the older V4 Flash benchmark result is not carried forward as if it measured V4.1. The Labs roster now targets the exact V4.1 Flash route; see the current leaderboard for the measured result and run timestamp.

The Price Comparison That Matters

Here's how DeepSeek stacks up against the competition:

ModelInput/1MOutput/1Mvs DeepSeek
DeepSeek V4.1 Flash$0.15$0.6—
GPT-5.6 Terra$2.00$12.0013x more expensive
Claude Opus 5.5$4.00$20.0027x more expensive
Gemini 3.1 Pro$2.00$12.0013x more expensive

Automatic Caching: The Hidden Advantage

Unlike providers that require explicit cache controls, DeepSeek's disk-based system caches repeated prefixes automatically. On V4.1 Flash, cache-hit input costs $0.003/M tokens off-peak, making stable agent prefixes unusually inexpensive.

Trade-offs to Consider

DeepSeek isn't perfect for every use case:

  • Rate limits can be tighter than OpenAI/Google during peak hours
  • Content filtering differs from Western providers
  • No free tier like Google Gemini offers
  • Fewer model options — just two models vs OpenAI's 13+

For most developers building cost-sensitive applications, these trade-offs are worth the 90% savings.

DeepSeek on Third-Party Hosts

You can also access DeepSeek models through inference providers like Groq, Together AI, and Fireworks — often with faster inference speeds but at higher prices. Check our full pricing comparison to compare.

Price History

Only models with a recorded price change are charted here.

DeepSeek V4.1 Flash

DeepSeek V4 Pro 0813

Price history tracking started April 2026. Flat model charts stay hidden until a price change is detected.
View pricing changelog →

Frequently asked questions

How much does DeepSeek cost per token?

DeepSeek V4.1 Flash costs $0.15 per 1M cache-miss input tokens and $0.6 per 1M output tokens off-peak. Peak rates are double. DeepSeek V4 Pro 0813 remains $0.66 input and $1.98 output off-peak.

What are DeepSeek peak hours?

DeepSeek peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC, Monday through Friday. All other hours, including weekends, use the half-price off-peak tier.

Does DeepSeek have a free tier?

DeepSeek does not offer a permanent free API tier. New accounts sometimes receive promotional balance, but production usage is pay-as-you-go.

How does DeepSeek compare to OpenAI?

DeepSeek V4.1 Flash at $0.15/$0.6 per 1M is about 13x cheaper than GPT-5.6 Terra on input and 20x cheaper on output. The trade-off is a smaller ecosystem and less battle-tested tooling than OpenAI.

What models does DeepSeek offer?

DeepSeek currently exposes DeepSeek V4.1 Flash on the new deepseek-flash route and DeepSeek V4 Pro 0813. The older deepseek-v4-flash and deepseek-v4-flash-vision-exp names temporarily route to V4.1 Flash.

Is DeepSeek V4 Pro being discontinued?

Not on the latest official pricing page. DeepSeek initially announced a September 14 redirect and repeated it by email, but its current pricing-page footnote says V4 Pro service will continue after September 14 with unchanged billing until further notice.

What's DeepSeek's context window?

DeepSeek lists a 1M-token context window and up to 384K max output for the current V4 family, which puts it in the long-context tier rather than the old 128K class.

Does DeepSeek have automatic caching?

Yes. DeepSeek applies cache-hit pricing automatically on repeated prefixes. On DeepSeek V4.1 Flash, cached input drops to $0.003 per 1M tokens off-peak.

Methodology

Pricing sourced from DeepSeek's pricing documentation and version details from its official changelog on . All prices in USD per 1 million tokens. Raw data: /api/pricing.json. API docs.

Compare All Providers

See how DeepSeek's pricing compares to every major AI API provider.

Further reading: DeepSeek vs ChatGPT pricing · Cheapest AI APIs in 2026.