Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
news

DeepSeek V4.1 Flash Pricing: V4 Pro Stays Live

DeepSeek V4.1 Flash cuts API prices and adds vision. See peak hours, endpoint migration, and why V4 Pro is no longer ending September 14.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • DeepSeek V4.1 Flash is live on the new deepseek-flash route with native image input, a 1M-token context window, and up to 384K output.
  • Off-peak rates are $0.15 cache-miss input, $0.003 cache-hit input, and $0.60 output per 1M tokens; weekday peak windows cost exactly twice as much.
  • The old deepseek-v4-flash and deepseek-v4-flash-vision-exp names temporarily route to V4.1 Flash, so teams should migrate to the new model name.
  • DeepSeek initially announced a September 14 V4 Pro redirect, but its latest pricing-page footnote now says Pro will continue with unchanged billing until further notice.

Current DeepSeek peak and off-peak tiers

Rates effective since September 10, 2026 at 4:00 AM UTC. Prices are USD per 1M tokens.

ModelRate periodInputCachedOutputInput / output vs current
DeepSeek V4.1 FlashOff-peak baseline$0.15$0.0030$0.6Comparison baseline
DeepSeek V4.1 FlashOff-peak$0.15$0.0030$0.6No change / No change
DeepSeek V4.1 FlashPeak$0.3$0.0060$1.20100% higher / 100% higher
DeepSeek V4 Pro 0813Off-peak baseline$0.66$0.022$1.98Comparison baseline
DeepSeek V4 Pro 0813Off-peak$0.66$0.022$1.98No change / No change
DeepSeek V4 Pro 0813Peak$1.32$0.044$3.96100% higher / 100% higher

Peak windows: 01:00-04:00 and 06:00-10:00 UTC; all other hours are off-peak. Built from the scheduled rate card in the canonical pricing API.

Cost comparison from today's pricing data

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$1.98DS V4.1 Flashdeepseek$0.15$0.6DS V4 Pro 0813deepseek$0.66$1.98

Estimate a DeepSeek workload

Assumes 75% input tokens and 25% output tokens using current per-million rates.

DeepSeek V4.1 Flash

deepseek

$2.63

Input share
$1.13
Output share
$1.50

DeepSeek V4 Pro 0813

deepseek

$9.90

Input share
$4.95
Output share
$4.95

V4.1 Flash price change

Model Previous input / output New input / output Change
DeepSeek V4.1 Flash $0.15 / $0.6 $0.15 / $0.6 No change

Previous rates come from the 2026-09-16 daily snapshot; new rates come from today's pricing dataset. Prices are per 1 million tokens.

Current DeepSeek API comparison rates

Model Provider Input / 1M Cached / 1M Output / 1M
DeepSeek V4.1 Flash deepseek $0.15 $0.0030 $0.6
DeepSeek V4 Pro 0813 deepseek $0.66 $0.022 $1.98

Built from pricing.json at publish time.

DeepSeek released V4.1 Flash on September 10, 2026 with a new API route, native visual understanding, and lower token prices. It also changed its V4 Pro transition plan after the launch announcement and customer email went out.

The buyer headline is simple: use deepseek-flash for new integrations, schedule flexible workloads outside the two weekday peak windows, and do not assume V4 Pro will disappear on September 14.

How much does DeepSeek V4.1 Flash cost?

DeepSeek’s current official rate card charges the following per million tokens:

Billing periodCache-hit inputCache-miss inputOutput
Off-peak$0.003$0.15$0.60
Peak$0.006$0.30$1.20

Peak hours are 01:00–04:00 UTC and 06:00–10:00 UTC, Monday through Friday. Every other period, including weekends, is off-peak. The live tables above are generated from our maintained pricing dataset and supersede these launch-day figures if DeepSeek changes the rate card.

Compared with the preceding V4 Flash schedule, V4.1 cuts cache-miss input by about 32%, cache-hit input by about 57%, and output by about 9% in both billing periods. For one million uncached input tokens plus one million output tokens, that is $0.75 off-peak or $1.50 at peak.

The endpoint migration matters

The current model name is deepseek-flash. DeepSeek says the older deepseek-v4-flash and deepseek-v4-flash-vision-exp names are still accepted temporarily, but requests are served by V4.1 Flash and billed at its price.

That makes the alias convenient for emergency compatibility, not a durable pin. Teams should move configuration, allowlists, dashboards, eval labels, and cost attribution to the new route. A request sent to the old name no longer proves that the old 0731 model produced the output.

V4.1 Flash is also a capability change. DeepSeek describes a 552B-parameter mixture-of-experts model with a new Causal Encoder–Decoder architecture, activating 8B parameters for input and 16B for output. The official pricing page lists a 1M-token context window, up to 384K output, and native vision support. Those specifications do not independently prove DeepSeek’s broader performance claims.

V4 Pro is staying—for now

DeepSeek’s September 10 announcement said deepseek-v4-pro would route to V4.1 Flash at 04:00 UTC on September 14. The email sent to API users repeated that transition, calling it a postponement of the earlier discontinuation date.

The current official pricing page now says something newer: in response to user demand, DeepSeek will continue providing V4 Pro after September 14 with unchanged billing and give further notice before any change. We therefore keep V4 Pro active in the DeepSeek pricing table rather than publishing a future redirect as settled fact.

This reversal is exactly why production teams should monitor both model behavior and billing. Pin regression tests, log the returned model identifier where available, and alert on official pricing-page changes instead of relying on a single launch email.

What buyers should do now

  1. Switch new workloads to deepseek-flash; treat legacy aliases as temporary compatibility routes.
  2. Re-run task-level evaluations because V4.1 is a different served model, not merely a price change.
  3. Split weekday peak and off-peak usage in cost dashboards. A monthly average can hide a 2x scheduling penalty.
  4. Keep V4 Pro tests running if its behavior matters, but prepare a migration plan because DeepSeek still promises only “further notice.”
  5. Compare cost per accepted task, including retries and latency, with the AI token calculator and Labs leaderboard.

For a managed third-party route, Novita also lists DeepSeek models; verify its current hosted model ID and rates in the live comparison before routing production traffic. Affiliate disclosure: this sponsored link may earn us a commission, without changing our pricing analysis or table order.

Labs treatment

The previous DeepSeek V4 Flash result cannot be relabeled as V4.1 evidence. A fresh September 11 run against the exact deepseek/deepseek-v4.1-flash route completed all 49 deterministic tasks correctly with zero API errors. The direct-list-price calculation was $0.003233 for the run; OpenRouter billed $0.006057. That narrow suite is useful for cost and regression tracking, not proof that V4.1 wins every workload. V4 Pro remains a distinct row while its first-party service continues. See the Labs coverage notes before comparing accuracy or cost-per-correct-task.

Sources: DeepSeek’s official V4.1 Flash release announcement, current models and pricing page, model repository, and Peter’s September 11 customer email. Official pages checked September 11, 2026.