Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
news · Updated July 29, 2026

Kimi K3 Pricing: API Discount vs Self-Hosting

Kimi K3 costs 10% less through Telnyx, while imec's self-host test finds better task resolution but higher hardware cost and slower results.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • Telnyx launched Kimi K3 on July 28 at 10% below Moonshot AI's direct input, cached-input, and output rates.
  • The route is OpenAI-compatible and includes automatic prefix caching, native vision, tool calling, and configurable reasoning.
  • imec's self-host test needed an 8×B300 node at about 20% higher hardware cost than GLM-5.2 on 8×B200, but resolved 86.4% versus 62.5% of tasks.
  • Existing Kimi users should benchmark managed and self-hosted routes on cost per accepted task before switching production traffic.

Kimi K3 and Claude Sonnet 5 workload cost

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$15.00Kimi K3telnyx$2.70$13.50Kimi K3moonshot$3.00$15.00Sonnet 5anthropic$2.00$10.00

Calculate Kimi K3 route costs

Assumes 75% input tokens and 25% output tokens using current per-million rates.

Claude Sonnet 5

anthropic

$40.00

Input share
$15.00
Output share
$25.00

Kimi K3

telnyx

$54.00

Input share
$20.25
Output share
$33.75

Kimi K3

moonshot

$60.00

Input share
$22.50
Output share
$37.50

Kimi K3 route prices and a premium-model reference

Model Provider Input / 1M Cached / 1M Output / 1M
Kimi K3 telnyx $2.70 $0.27 $13.50
Kimi K3 moonshot $3.00 $0.3 $15.00
Kimi K3 novita $3.00 $0.3 $15.00
Claude Sonnet 5 anthropic $2.00 $0.2 $10.00

Built from pricing.json at publish time.

Telnyx added Kimi K3 to its Inference API on July 28, creating a new managed route for Moonshot AI’s flagship open-weight model. The launch matters because Telnyx prices all three token categories 10% below Moonshot’s direct rate. The live table above pulls the exact rates from our pricing dataset at build time.

The new endpoint uses the model ID moonshotai/Kimi-K3 and an OpenAI-compatible API. Telnyx says it runs the model on company-owned GPU infrastructure rather than forwarding requests to another model host.

What changed

Telnyx now offers Kimi K3 with its one-million-token context window, native image and video input, tool calling, JSON-schema output, dynamic tool loading, and three configurable reasoning levels. Automatic prefix caching is enabled for repeated prompt prefixes.

This is a hosting and pricing event, not a new Kimi model release. Moonshot launched K3 earlier in July; today’s change adds a cheaper production route and another infrastructure vendor to evaluate.

The 10% discount applies consistently to fresh input, cached input, and output. That makes the comparison simple, but token price alone does not settle the buying decision. Rate limits, latency, cache behavior, regional availability, support, and data-handling terms can outweigh a modest rate difference.

Self-hosting update: K3 needs B300

imec’s July 29 update adds a very different Kimi K3 route. The 1.4TB model did not leave enough KV-cache headroom on its 8×B200 system, so the team moved to an 8×B300 node with 2.3TB of total HBM. It estimates that hardware at roughly 20% more than the B200 setup used for GLM-5.2.

imec resultKimi K3GLM-5.2
Concurrent sessions1624
Aggregate output at 16 users122 tok/s170 tok/s
Median task time38 min26 min
Tasks resolved86.4%62.5%

K3 traded about 30% lower throughput and 50% longer task time for a 24-point task-resolution lead. The quality result needs caution: imec says its SWE-Bench Pro subset may have appeared in K3’s training data, and only the article’s graphs were updated for the new run.

The headline’s “20% better task resolution” wording understates the published table: 86.4% versus 62.5% is a 23.9 percentage-point lead, or roughly 38% higher relative resolution. It is still one potentially contaminated 64-task subset, not a guaranteed production uplift.

Labs status

Kimi K3 has a verified moonshotai/kimi-k3 route and is queued for the AI Pricing Guru Labs leaderboard, but it is not yet in the live published result. The July 29 full-roster refresh exhausted the benchmark account’s available OpenRouter credit before all model-task pairs completed. The publication guard rejected that partial run, so we will not present an incomplete score.

This is the explicit availability blocker until a complete 49-task Kimi run can be funded and rerun. Even after inclusion, Labs will compare public endpoint cost per correct answer; it will not reproduce imec’s 8×B300 deployment, SGLang serving stack, concurrency test, or hardware economics.

What this means for API buyers

Teams already testing Kimi K3 gain immediate negotiating and routing leverage. A provider-neutral gateway can send eligible traffic to Telnyx while retaining Moonshot or Novita as fallback routes. The model weights may be the same, but serving behavior can differ enough to affect real cost per completed task.

Long-context coding and agent workflows stand to benefit most. Stable system prompts, repository summaries, and tool schemas can reuse cached prefixes, while dynamic tool loading can keep irrelevant schemas out of the prompt. Both features reduce billed tokens when implemented carefully.

Teams comparing Kimi with Claude should use the Anthropic pricing page and the chart above as a starting point, then test identical tasks. For lower-cost open-model routing, compare the Together AI pricing page and our Z.ai vs DeepSeek cost guide.

Teams that want Kimi K3 without operating a B300 node can benchmark Novita’s managed Kimi K3 route alongside Telnyx and Moonshot.

Affiliate disclosure: we may earn a commission from the sponsored link above. It does not affect our pricing analysis.

Who should switch—and who should wait

Switch a small traffic slice if Kimi K3 already passes your quality bar and your workload has meaningful output or repeated context. Measure first-token latency, total completion time, cache-hit rate, tool-call success, retries, and accepted outputs.

Wait if you need firm regional controls, high committed throughput, or mature enterprise support. Confirm those operational details with Telnyx before moving sensitive or business-critical workloads.

Do not move routine extraction, classification, or short support replies to Kimi K3 solely because the route is discounted. Smaller models can still deliver a lower cost per successful task.

Practical next step

Run the same production-shaped evaluation through Telnyx and your current provider. Use the inline calculator above or the full AI token calculator to model volume, but make the final decision on cost per accepted result rather than price per million tokens.

Start with a limited canary, keep a fallback route, and compare invoices with request-level telemetry. If Telnyx preserves quality and reliability, its uniform discount becomes a straightforward saving. If retries or latency rise, the cheaper token rate can disappear quickly.

Sources: Telnyx Kimi K3 release note, Moonshot Kimi K3 pricing, imec’s Kimi K3 self-hosting update, and the live AI Pricing Guru API dataset.