Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
news

DeepSeek API Pricing Update — Impact (August 2026)

DeepSeek introduces peak and off-peak API rates on August 16. Even off-peak V4 pricing rises versus today's rate card.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • DeepSeek will switch V4 Flash and V4 Pro to peak and off-peak billing at 16:00 UTC on August 16, 2026.
  • Off-peak is 50% below the new peak rate, but it is not a discount from today's price: every listed V4 token rate still rises off-peak.
  • Batch flexible jobs outside the two peak windows; latency-sensitive teams should budget for the peak rate before the cutover.

DeepSeek V4 pricing: current vs August 16 rates

New rates take effect August 16, 2026 at 4:00 PM UTC. Prices are USD per 1M tokens.

ModelRate periodInputCachedOutputInput / output vs current
DeepSeek V4 Flash 0731Current$0.14$0.0028$0.28Baseline
DeepSeek V4 Flash 0731Off-peak$0.22$0.0070$0.6657% higher / 136% higher
DeepSeek V4 Flash 0731Peak$0.44$0.014$1.32214% higher / 371% higher
DeepSeek V4 Pro 0813Current$0.435$0.0036$0.87Baseline
DeepSeek V4 Pro 0813Off-peak$0.66$0.022$1.9852% higher / 128% higher
DeepSeek V4 Pro 0813Peak$1.32$0.044$3.96203% higher / 355% higher

Peak windows: 01:00-04:00 and 06:00-10:00 UTC; all other hours are off-peak. Built from the scheduled rate card in the canonical pricing API.

Cost comparison from today's pricing data

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$4.40DS V4 Flash 0731deepseek$0.14$0.28DS V4 Pro 0813deepseek$0.435$0.87GLM-5.2zai$1.40$4.40GPT 5.6 Lunaopenai$0.2$1.20

Estimate a workload at currently active rates

Assumes 75% input tokens and 25% output tokens using current per-million rates.

DeepSeek V4 Flash 0731

deepseek

$1.75

Input share
$1.05
Output share
$0.70

GPT-5.6 Luna

openai

$4.50

Input share
$1.50
Output share
$3.00

DeepSeek V4 Pro 0813

deepseek

$5.44

Input share
$3.26
Output share
$2.18

GLM-5.2

zai

$21.50

Input share
$10.50
Output share
$11.00

Current rates before the August 16 cutover

Model Provider Input / 1M Cached / 1M Output / 1M
DeepSeek V4 Flash 0731 deepseek $0.14 $0.0028 $0.28
DeepSeek V4 Pro 0813 deepseek $0.435 $0.0036 $0.87
GLM-5.2 zai $1.40 $0.26 $4.40
GPT-5.6 Luna openai $0.2 $0.02 $1.20

Built from pricing.json at publish time.

DeepSeek announced a new peak and off-peak API rate card for V4 Flash and V4 Pro. It takes effect at 16:00 UTC on August 16, 2026.

The headline needs careful reading. DeepSeek says off-peak rates are 50% below peak rates, but the scheduled table above shows that off-peak is still more expensive than the currently active rate in every input, cache-hit, and output category. This is a price increase with a time-of-day discount—not a general 50% price cut.

Use the live DeepSeek pricing page for the active rate card, the token calculator for workload estimates, and our DeepSeek API pricing guide for model-selection context.

What changes on August 16

The V4 lineup moves from one always-on rate to two billing periods. Peak windows run from 01:00–04:00 UTC and 06:00–10:00 UTC. All other hours are off-peak.

That creates 17 off-peak hours and seven peak hours per UTC day. Workload timing now becomes a first-order cost variable alongside model choice, token volume, cache hits, and output length.

The generated table above reads both the active and scheduled values from our canonical pricing dataset. DeepSeek’s current rate remains active until the announced cutover.

Who benefits—and who pays more

Batch users gain the most control. Document extraction, offline evaluation, data enrichment, synthetic-data generation, nightly coding checks, and queue-based support processing can be scheduled outside the two peak windows.

Global interactive products have less room to maneuver. A customer-facing agent cannot always delay a response because traffic lands during peak hours. Those teams should budget against the peak tier and treat any off-peak execution as savings, rather than building forecasts around the lower tier.

Cache-heavy agent systems face an especially important recheck. The new schedule changes cache-hit rates as well as uncached input and output. Replaying only short prompts will miss the impact on long repeated system prompts, tool schemas, repository maps, and conversation prefixes.

What this means for DeepSeek’s value case

DeepSeek V4 remains inexpensive relative to many premium frontier APIs, but its cost advantage narrows after the cutover. The decision should now compare accepted-task cost under the actual traffic schedule—not a single advertised rate.

V4 Flash is still the natural starting point for high-volume, easily checked work. V4 Pro belongs on harder coding, reasoning, and analysis tasks where fewer retries justify the higher unit cost. Compare both against current OpenAI API pricing and the build-time table above before changing a production route.

For a managed alternative with multiple open-model routes, check Novita’s current catalog and benchmark the same accepted tasks, latency, and retry rate.

Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored link above at no extra cost to you. It does not affect our analysis.

What DeepSeek API users should do now

  1. Export hourly token usage in UTC and map it against both peak windows.
  2. Move delay-tolerant queues to off-peak hours before the August 16 cutover.
  3. Budget synchronous traffic at peak rates; do not assume every request can capture the discount.
  4. Track cached and uncached input separately, because both categories change.
  5. Re-run model evaluations on V4 Flash and V4 Pro and compare cost per accepted result.
  6. Add a spend alert for the first full billing day after the new schedule starts.

Bottom line

DeepSeek’s new time-of-day billing rewards schedulable workloads, but it raises the V4 rate card even during off-peak hours. Teams that can shift batch jobs gain a meaningful discount from the new peak tier; teams serving live traffic should plan for the higher peak bill.

Sources: DeepSeek’s official pricing documentation and its August 13 pricing announcement. Current and scheduled machine-readable rates are available in AI Pricing Guru’s pricing API.