Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
news

DeepSeek V4.1 Flash Runs on a 16GB M1 Mac

DeepSeek V4.1 Flash ran on a 2020 M1 Mac mini via SSD streaming—but at 23 seconds per token. See the practical cost and latency impact.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • An independent experiment ran the original 475 GiB DeepSeek V4.1 Flash FP4/FP8 checkpoint on a 2020 M1 Mac mini with 16 GB unified memory by streaming selected weights from its internal SSD.
  • The optimized run took about 108 seconds to produce the first token and 22.8 seconds per subsequent token—roughly 0.044 tokens per second.
  • This is a real systems-engineering result, but not an interactive local-LLM experience: a 2,048-token response would take about 13 hours if the short-run rate held.
  • DeepSeek's API prices did not change. The hosted V4.1 Flash route remains the practical option for normal interactive and production work.

Hosted DeepSeek V4.1 Flash versus V4 Pro

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$1.98DS V4.1 Flashdeepseek$0.15$0.6DS V4 Pro 0813deepseek$0.66$1.98

Estimate a hosted DeepSeek workload

Assumes 75% input tokens and 25% output tokens using current per-million rates.

DeepSeek V4.1 Flash

deepseek

$2.63

Input share
$1.13
Output share
$1.50

DeepSeek V4 Pro 0813

deepseek

$9.90

Input share
$4.95
Output share
$4.95

Current hosted DeepSeek V4-family rates

Model Provider Input / 1M Cached / 1M Output / 1M
DeepSeek V4.1 Flash deepseek $0.15 $0.0030 $0.6
DeepSeek V4 Pro 0813 deepseek $0.66 $0.022 $1.98

Built from pricing.json at publish time.

DeepSeek V4.1 Flash has run locally on a 2020 M1 Mac mini with 16 GB of unified memory. The catch is in the unit: the optimized experiment generated one token every 22.8 seconds, not 22.8 tokens per second.

The independent project streamed the official checkpoint from SSD through a custom MLX runner. It is a striking proof that a model far larger than system memory can execute on entry-level Apple Silicon, but it does not make a five-year-old Mac a practical substitute for hosted inference.

What the experiment achieved

The test used the original 475 GiB of FP4/FP8 weights—about 510 GB in decimal units—across 48 safetensors shards. There was no additional quantization and no smaller replacement model. The machine had an eight-core CPU, eight-core GPU, 16 GB of unified memory, and an internal 1 TB SSD.

Instead of loading the whole checkpoint, the runner read only the six routed experts selected for each token, fetched required rows from the model’s Engram tables, retained a bounded dense-weight cache, and preserved the KV state between tokens. DeepSeek’s model card says V4.1 Flash has a 552-billion-parameter backbone but activates 16 billion parameters during decode, making sparse loading possible in principle.

Measured configurationTime to first tokenDecode speedActive MLX memory
Original SSD streaming128.3 seconds30.7 sec/token1.67 GiB
Reused buffers + 4 GiB cache + compiled decoding108.2 seconds22.6 sec/token5.65 GiB
Optimized eight-token repeat108.3 seconds22.8 sec/token5.65 GiB

The repeat was about 1.35 times faster than the fresh baseline. These were short exploratory runs, not randomized trials: the OS file cache was not cleared, other apps remained open, and active MLX memory did not include all system memory or swap.

Why 23 seconds per token matters

At 22.8 seconds per token, throughput is about 0.044 tokens per second, or 158 tokens per hour. The repository estimates that a 2,048-token allowance would take about 13 hours if the short-run speed held. A brief eight-token test still takes roughly four and a half minutes once first-token latency is included.

That makes the setup unsuitable for chat, coding agents, batch production, or any workflow with retries. It is valuable as a demonstration of bounded-memory inference and as a test bed for prefetching, caching, expert routing, and SSD-aware runtimes.

Did DeepSeek API pricing change?

No. DeepSeek’s official rate card still lists V4.1 Flash at $0.003 cache-hit input, $0.15 cache-miss input, and $0.60 output per million tokens off-peak. The two documented weekday peak windows cost exactly twice as much. The live table and chart above come from our daily-maintained pricing dataset and supersede those checked figures if DeepSeek changes the rate card.

The local route also requires more than the Mac’s purchase price: at least 550–600 GB of free fast SSD space, a roughly 510 GB download, setup time, electricity, storage wear, and hours of runtime per answer. Hosted inference bills tokens and returns them at serving speed. A Hacker News commenter calculated that a four-million-token task would take about 2.9 years at the Mac experiment’s measured decode rate.

For most developers, the decision is not close. API inference wins on latency, concurrency, and opportunity cost. Local execution wins only when the experiment itself is the goal, data cannot leave the device, or offline access matters more than completion time. Compare that conclusion with our earlier DeepSeek hosted-versus-local break-even analysis.

Labs inclusion and the local-hardware blocker

DeepSeek V4.1 Flash is already included in AI Pricing Guru Labs. Its fresh managed-route run attempted all 49 deterministic text tasks, answered all 49 correctly, and returned no endpoint errors. That result measures cost per correct answer through a public endpoint.

The M1 result remains a separate local-throughput coverage note. A fair comparison needs the repository and weight revisions pinned, repeat runs on independently controlled hardware, identical prompts and output lengths across local and hosted routes, warm and cold latency, energy use, failure rates, and a full hardware-cost ledger. Our current text leaderboard cannot turn a storage-specific seconds-per-token result into a comparable API score.

What DeepSeek users should do

  1. Use deepseek-flash or a verified hosted V4.1 Flash route for interactive and production work.
  2. Measure end-to-end cost per accepted task, including output length, retries, latency, and human waiting time.
  3. Reproduce the Mac setup only with a fast SSD, ample free space, and workloads that can tolerate overnight generation.
  4. Treat the published speed as a project-specific result until it is repeated across machines, prompts, and longer outputs.

For a managed third-party route, compare Novita’s listed DeepSeek V4.1 Flash route with the first-party API, confirming the exact model and billing tier first. You can also compare current Novita pricing and OpenAI pricing as latency and cost controls.

Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored Novita link at no extra cost to you. It does not affect this analysis.

Bottom line

DeepSeek V4.1 Flash genuinely ran on a 16 GB M1 Mac mini without replacing the official weights. The achievement is the streaming system, not the user experience. At roughly 0.044 tokens per second, this is research-grade local inference; the hosted API remains the practical choice.

Sources: FP4 Brain’s September 11 post, the project’s reproducible GitHub repository, DeepSeek’s official V4.1 Flash model card, official API pricing, and the Hacker News discussion, checked September 12, 2026.