Moonshot AI Kimi K3 Pricing
Last updated:
Kimi K3 costs $3.00 input / $15.00 output per 1M tokens through Moonshot AI. Cached input is $0.30 per 1M tokens. The official endpoint has a 1,048,576-token context window, while the open-weight release gives teams a separate self-hosting path with a very different infrastructure bill.
Kimi K3 API token prices
| Try it | |||||||
|---|---|---|---|---|---|---|---|
Kimi K3 | Moonshot | Mid | 1M | $3.00 | $0.30 | $15.00 | Try API → |
- Kimi K3MoonshotMid
- Input
- $3.00
- Cached
- $0.30
- Output
- $15.00
Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.
All product names, logos, and brands are property of their respective owners and are used for identification purposes only.
These are Moonshot AI's first-party list prices. Hosted providers can publish different rates, latency, caching, or support terms for the same model family.
Managed API versus self-hosting
Moonshot's managed API is the simplest route for variable traffic: you pay for tokens and avoid owning an accelerator pool. Self-hosting exchanges that bill for fixed hardware capacity, serving engineering, monitoring, upgrades, and utilization risk.
The July 29 imec test makes the scale concrete. Its roughly 1.4TB Kimi K3 checkpoint did not fit on an 8×B200 node with enough room left for KV cache. The team moved to 8×B300, estimated the hardware at about 20% more than its GLM-5.2 B200 setup, and measured lower throughput and longer task time—but a 23.9-point task-resolution lead on its 64-task subset.
Read the complete methodology and contamination caveat in our Kimi K3 API versus self-hosting analysis. Use the token cost calculator to price managed volume before comparing it with a continuously rented B300 node.
Open weights do not mean zero cost
Moonshot describes Kimi K3 as a 2.8T-parameter mixture-of-experts model with 104B active parameters, MXFP4 weights, native multimodality, and a one-million-token context window. The weights are downloadable under the Kimi K3 License, but teams still need to review those terms and budget for storage, HBM, KV cache, power or rental, networking, and on-call operations.
If operating a multi-GPU node is not economical, compare Novita's managed Kimi K3 route with Moonshot's first-party API and other hosts.
Affiliate disclosure: we may earn a commission from the sponsored link above. It does not affect the pricing table or analysis.
What the Kimi K3 “LLM City” shows
The independent LLM City visualization maps Kimi K3's tensor shapes into a 3D architecture and an exact-area atlas. It counts 2,779,931,834,976 logical weight positions—consistent with Moonshot's rounded 2.8T total—and assigns every scalar a user-selectable physical tile. At the default 2.5 mm setting, that area is about 17.37 km², equal to a square roughly 4.17 km on each side.
This is an architecture-scale metaphor, not a dump of learned values. The page says its close-up cell numbers are deterministic mock values. Also keep total and active parameters separate: Moonshot says K3 activates about 104B parameters per token, so the 2.8T city is useful for understanding storage and deployment scale, not as a direct measure of per-token compute or API price.
Read our Kimi K3 model-size and LLM City cost analysis for the scale math and buyer implications.
What the four-SSD MacBook result proves
The independent ARGODRIVE Deltafin project reports 1.0015 tokens/s of steady decode for a 512-token raw completion on an M5 Max MacBook Pro with 128GB unified memory. The setup streamed Kimi K3's expert bank from the internal SSD and three external NVMe enclosures. The same package reports about 375 seconds to the first token for a 512-token prompt.
This is a reproducible systems experiment, not a Moonshot performance claim or a price change. Its four-drive ladder, project-specific storage layout, speculative drafter, int8 resident spine, cold-run method, and prompt dependence make the result unsuitable as a universal local-speed estimate. Read the Kimi K3 MacBook Pro cost and latency analysis for the measurement boundaries and managed-API comparison.
Labs status
Kimi K3 is now live in the AI Pricing Guru Labs leaderboard. The complete August 17 run attempted all 49 deterministic tasks with no endpoint errors and scored 47 correct. The direct-price estimate was $0.135252 for the run, while OpenRouter billed $0.13395165.
Labs still does not test the private imec task subset, SGLang configuration, B300 utilization, Deltafin's four-SSD local throughput, vision performance, or LLM City's architectural representation. It measures public-endpoint cost per correct answer on a small text suite. The local replay requirements are documented in our benchmark coverage notes.
Frequently asked questions
How much does the Kimi K3 API cost?
Moonshot AI lists Kimi K3 at $3.00 per 1M uncached input tokens, $0.30 per 1M cached input tokens, and $15.00 per 1M output tokens.
Can Kimi K3 be self-hosted?
Yes. Moonshot AI publishes Kimi K3 as an open-weight 2.8-trillion-parameter model under the Kimi K3 License. The model card lists MXFP4 weights and 104 billion activated parameters, so production hosting still requires substantial accelerator memory.
What hardware does Kimi K3 need?
Moonshot AI does not publish one universal minimum server. In imec’s July 2026 test, roughly 1.4TB of weights did not leave KV-cache headroom on 8×B200, so the team used 8×B300 with 2.3TB total HBM. Quantization, context, concurrency, and serving software change the requirement.
Can Kimi K3 run on a MacBook Pro?
An independent Deltafin benchmark reports about 1 token per second of steady decode on an M5 Max MacBook Pro with 128GB unified memory and four SSDs. It also reports roughly 375 seconds to the first token for a 512-token prompt, so this is an experimental local route rather than an API-speed replacement.
Is Kimi K3 included in AI Pricing Guru Labs?
Yes. The August 17 live roster includes a complete 49-task Kimi K3 run: 47 correct, no endpoint errors, and $0.13395 billed through OpenRouter. Labs measures a small deterministic text suite, not self-hosting throughput or the LLM City visualization.