Moonshot AI Kimi K3 Pricing
Last updated:
Kimi K3 costs $3.00 input / $15.00 output per 1M tokens through Moonshot AI. Cached input is $0.30 per 1M tokens. The official endpoint has a 1,048,576-token context window, while the open-weight release gives teams a separate self-hosting path with a very different infrastructure bill.
Kimi K3 API token prices
Kimi K3 | Moonshot | Mid | 1M | $3.00 | $0.30 | $15.00 |
- Kimi K3MoonshotMid
- Input
- $3.00
- Cached
- $0.30
- Output
- $15.00
Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.
All product names, logos, and brands are property of their respective owners and are used for identification purposes only.
These are Moonshot AI's first-party list prices. Hosted providers can publish different rates, latency, caching, or support terms for the same model family.
Managed API versus self-hosting
Moonshot's managed API is the simplest route for variable traffic: you pay for tokens and avoid owning an accelerator pool. Self-hosting exchanges that bill for fixed hardware capacity, serving engineering, monitoring, upgrades, and utilization risk.
The July 29 imec test makes the scale concrete. Its roughly 1.4TB Kimi K3 checkpoint did not fit on an 8×B200 node with enough room left for KV cache. The team moved to 8×B300, estimated the hardware at about 20% more than its GLM-5.2 B200 setup, and measured lower throughput and longer task time—but a 23.9-point task-resolution lead on its 64-task subset.
Read the complete methodology and contamination caveat in our Kimi K3 API versus self-hosting analysis. Use the token cost calculator to price managed volume before comparing it with a continuously rented B300 node.
Open weights do not mean zero cost
Moonshot describes Kimi K3 as a 2.8T-parameter mixture-of-experts model with 104B active parameters, MXFP4 weights, native multimodality, and a one-million-token context window. The weights are downloadable under the Kimi K3 License, but teams still need to review those terms and budget for storage, HBM, KV cache, power or rental, networking, and on-call operations.
If operating a multi-GPU node is not economical, compare Novita's managed Kimi K3 route with Moonshot's first-party API and other hosts.
Affiliate disclosure: we may earn a commission from the sponsored link above. It does not affect the pricing table or analysis.
Labs status
Kimi K3 is not yet in the live AI Pricing Guru Labs leaderboard. Its moonshotai/kimi-k3 route is verified and queued, but the July 29 full-roster refresh exhausted the benchmark account's available OpenRouter credit before all model-task pairs completed. The publication guard rejected the partial result. A complete 49-task rerun is the explicit inclusion blocker.
Even after inclusion, Labs will not test the private imec task subset, SGLang configuration, B300 utilization, or concurrency economics, so it cannot validate the self-hosting percentages by itself.
Frequently asked questions
How much does the Kimi K3 API cost?
Moonshot AI lists Kimi K3 at $3.00 per 1M uncached input tokens, $0.30 per 1M cached input tokens, and $15.00 per 1M output tokens.
Can Kimi K3 be self-hosted?
Yes. Moonshot AI publishes Kimi K3 as an open-weight 2.8-trillion-parameter model under the Kimi K3 License. The model card lists MXFP4 weights and 104 billion activated parameters, so production hosting still requires substantial accelerator memory.
What hardware does Kimi K3 need?
Moonshot AI does not publish one universal minimum server. In imec’s July 2026 test, roughly 1.4TB of weights did not leave KV-cache headroom on 8×B200, so the team used 8×B300 with 2.3TB total HBM. Quantization, context, concurrency, and serving software change the requirement.
Is Kimi K3 included in AI Pricing Guru Labs?
Not yet in the live published result. The public route is verified, but the July 29 full-roster refresh exhausted the benchmark account’s OpenRouter credit before all tasks completed. The publication guard rejected the partial run.