DeepSeek V4.1 Flash Hacking Benchmark: Cost & Limits
DeepSeek V4.1 Flash scored 11/11 in Enclave's cyber test for $4.65. See cache economics, audit limits, current API prices, and buyer guidance.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- Verdict: DeepSeek V4.1 Flash is a compelling low-cost cyber-agent candidate, but Enclave's 11/11 result is one private, source-assisted harness—not proof that it is universally the best hacking model.
- Enclave reports code execution on all 11 vulnerable targets, with all four fixed controls secure; six runs followed the intended exploit and five found alternate paths in the test environment.
- Accepted runs cost $4.65 and the complete run set cost $5.14. More than 99% of the reported input tokens were cache hits, so the headline is not a cold-prompt cost.
- No DeepSeek model, endpoint, or price changed. V4.1 Flash stays in Labs for its existing 49-task text result; the cyber claim needs a separate controlled security harness.
DeepSeek peak and off-peak pricing
Rates effective since September 10, 2026 at 4:00 AM UTC. Prices are USD per 1M tokens.
| Model | Rate period | Input | Cached | Output | Input / output vs current |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | Off-peak baseline | $0.15 | $0.0030 | $0.6 | Comparison baseline |
| DeepSeek V4.1 Flash | Off-peak | $0.15 | $0.0030 | $0.6 | No change / No change |
| DeepSeek V4.1 Flash | Peak | $0.3 | $0.0060 | $1.20 | 100% higher / 100% higher |
| DeepSeek V4 Pro 0813 | Off-peak baseline | $0.66 | $0.022 | $1.98 | Comparison baseline |
| DeepSeek V4 Pro 0813 | Off-peak | $0.66 | $0.022 | $1.98 | No change / No change |
| DeepSeek V4 Pro 0813 | Peak | $1.32 | $0.044 | $3.96 | 100% higher / 100% higher |
Peak windows: 01:00-04:00 and 06:00-10:00 UTC; all other hours are off-peak. Built from the scheduled rate card in the canonical pricing API.
Current DeepSeek API price scale
USD per 1M tokens. Input and output rates are charted separately.
Estimate your own DeepSeek agent workload
Assumes 75% input tokens and 25% output tokens using current per-million rates.
DeepSeek V4.1 Flash
deepseek
$2.63
- Input share
- $1.13
- Output share
- $1.50
DeepSeek V4 Pro 0813
deepseek
$9.90
- Input share
- $4.95
- Output share
- $4.95
Current DeepSeek rates—the benchmark did not change API pricing
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| DeepSeek V4.1 Flash | deepseek | $0.15 | $0.0030 | $0.6 |
| DeepSeek V4 Pro 0813 | deepseek | $0.66 | $0.022 | $1.98 |
Built from pricing.json at publish time.
DeepSeek V4.1 Flash just posted an unusually strong cyber-agent result: 11 verified code-execution outcomes across 11 vulnerable targets, while all four fixed controls stayed secure. Enclave says the accepted runs cost $4.65.
That is attention-worthy, but the useful conclusion is narrower than “best hacking model.” This was Enclave’s private benchmark against isolated Grafana, Jenkins, and Nextcloud environments. A path-level audit found that six successful runs used the intended weakness and five used other routes exposed by the vulnerable test versions.
No new DeepSeek model, endpoint, or rate launched with the result. The live tables above render the maintained DeepSeek API pricing and supersede any static price figure if the official rate card changes.
What Enclave’s 11/11 result actually measured
The benchmark gave V4.1 Flash source code and isolated vulnerable and fixed application versions. The agent could start services, send requests, inspect behavior, run shell commands, and change its approach after failures.
Enclave reported these aggregate results:
| Enclave-reported measure | Result |
|---|---|
| Vulnerable targets with verified code execution | 11 of 11 |
| Fixed controls that remained secure | 4 of 4 |
| Runs using the planned exploit path | 6 |
| Runs using alternate test-environment paths | 5 |
| Bash commands issued | 2,349 |
| Active model time | 2h 38m |
| Median successful run | 4m 38s |
| Accepted-run cost | $4.65 |
| Full cost including failed and replacement runs | $5.14 |
The planned-path wins included three Jenkins credential attacks, one Jenkins upload race, and two Nextcloud access-control attacks. The alternate paths were three Grafana temporary-plugin executions and two shorter Jenkins file-link routes.
Those alternate routes are not automatically a failure. Real agents are supposed to find working paths that evaluators did not anticipate. They do, however, change what the perfect score means: the original outcome grader verified code execution, while the later audit separated intended exploits from unintended benchmark shortcuts.
Enclave says the five alternate paths belonged to its private benchmark environment and do not claim new vulnerabilities in upstream Grafana or Jenkins. It has closed those extra routes and versioned the repaired challenges, so earlier model results need reruns before they can enter the updated comparison.
Why the $4.65 cost needs a cache label
Enclave reports 268.3 million input tokens and roughly two million output tokens. Of that input, 266.2 million tokens were cached—about 99.2% of all input.
That makes $4.65 a valid bill for this run, not a universal price for 11 cyber tasks. A cold run, a different source tree, lower prefix reuse, another time window, extra retries, or a different host’s cache accounting could cost materially more. The complete run set already moved from $4.65 to $5.14 once failed attempts and replacements were included.
DeepSeek’s official V4.1 Flash rate card remains unchanged. It bills cache-hit input, cache-miss input, and output separately, with two weekday peak windows at twice the off-peak rates. Use the live schedule above and the token cost calculator with your measured cache-hit ratio rather than multiplying Enclave’s total by an expected task count.
The buyer decision: evaluate cost per accepted security outcome
V4.1 Flash now deserves a place in a cyber-agent evaluation shortlist. The result combines low billed cost, long-horizon tool use, source reading, exploit adaptation, and correct restraint on the four fixed controls.
It does not justify autonomous production exploitation. A serious buyer test should track:
- accepted findings or verified target outcomes, not self-reported success;
- intended-path success and alternate-path success separately;
- fixed controls, false positives, destructive actions, and policy violations;
- cold and warm cache cost, retries, wall-clock time, and tool-infrastructure spend;
- exact application revisions, model route, prompts, tools, and stopping rules.
The strongest practical pattern is an isolated lab with explicit scope, deterministic proof checks, and human authorization before any action outside the sandbox. For a managed third-party route, compare Novita’s current DeepSeek listing with the first-party API, confirming the exact model ID, cache treatment, and rates first.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored link above at no extra cost to you. It does not affect this analysis.
Labs decision: include the model, block the claim
DeepSeek V4.1 Flash is already included in AI Pricing Guru Labs through the exact deepseek/deepseek-v4.1-flash managed route. Its current 49-task deterministic text run completed 49 of 49 tasks with no API errors.
We are not importing Enclave’s 11/11 score into that leaderboard and are not triggering another paid text rerun. The cyber test measures source-assisted agent behavior against instrumented applications; our leaderboard measures short deterministic text tasks and direct cost per correct answer.
The Labs coverage note records the blocker. A comparable cyber replay needs the repaired challenge versions, immutable container images, prompts and tool definitions, raw transcripts, outcome and path graders, token/cache ledgers, retry accounting, fixed controls, and the same model route. Until then, 11/11 is Enclave’s audited result—not ours.
Bottom line
DeepSeek V4.1 Flash delivered a strong and economically interesting cyber-agent result. The model achieved all 11 target outcomes, respected every fixed control, and exposed weaknesses in the benchmark itself for a small reported bill.
The honest buying takeaway is to test V4.1 Flash, not crown it universally. Preserve Enclave’s distinctions: six intended exploits, five benchmark-specific alternate routes, a cache-heavy $4.65 accepted-run bill, and a repaired harness that resets cross-model comparability.
Sources: Enclave’s DeepSeek V4.1 Flash hacking benchmark and path audit, DeepSeek’s official models and pricing page, V4.1 Flash announcement, and model repository. Benchmark claims, official pricing, route status, and linked sources checked September 16, 2026 at 16:44 UTC.