Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
news · Updated July 30, 2026

OpenAI GPT-5.6 Price Cut: Impact & What It Means

OpenAI cut GPT-5.6 Luna API prices by 80% and Terra by 20%. Compare old and new rates, Fast mode, and the best tier for each workload.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • Price winner: GPT-5.6 Luna. OpenAI cut its standard input, cached-input, and output rates by 80%.
  • GPT-5.6 Terra is 20% cheaper; GPT-5.6 Sol standard pricing is unchanged.
  • Fast mode replaces Priority Processing and offers up to 2.5x faster Sol responses at 2x the standard rate.
  • Use Luna for high-volume defined work, Terra for balanced production tasks, and Sol only where maximum intelligence or latency changes the outcome.

GPT-5.6 price-performance tiers vs Claude

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$50.00GPT 5.6 Lunaopenai$0.2$1.20GPT 5.6 Terraopenai$2.00$12.00GPT 5.6 Solopenai$5.00$30.00Sonnet 5anthropic$2.00$10.00Fable 5anthropic$10.00$50.00

Calculate your GPT-5.6 workload cost

Assumes 75% input tokens and 25% output tokens using current per-million rates.

GPT-5.6 Luna

openai

$4.50

Input share
$1.50
Output share
$3.00

Claude Sonnet 5

anthropic

$40.00

Input share
$15.00
Output share
$25.00

GPT-5.6 Terra

openai

$45.00

Input share
$15.00
Output share
$30.00

GPT-5.6 Sol

openai

$112.50

Input share
$37.50
Output share
$75.00

GPT-5.6 old vs new API pricing

Model Previous input / output New input / output Change
GPT-5.6 Luna $0.2 / $1.20 $0.2 / $1.20 No change
GPT-5.6 Terra $2.00 / $12.00 $2.00 / $12.00 No change
GPT-5.6 Sol $5.00 / $30.00 $5.00 / $30.00 No change

Previous rates come from the 2026-08-09 daily snapshot; new rates come from today's pricing dataset. Prices are per 1 million tokens.

OpenAI has passed part of GPT-5.6’s serving-efficiency gains to customers one day after detailing the engineering behind them. Starting July 30, Luna’s standard API rate is 80% lower and Terra’s is 20% lower. Sol’s standard rate is unchanged, but a new Fast mode offers higher speed for latency-sensitive work.

The chart, calculator, and old-vs-new table above are generated from our current and historical pricing datasets. See the complete lineup on the OpenAI pricing page or model your own volume in the AI token calculator.

What changed beyond the price cut

Cached-input reads remain 90% below standard input, while cache writes cost 1.25 times standard input. The largest absolute savings therefore go to high-volume work that already meets its quality target on Luna or Terra.

OpenAI says the lower Luna and Terra rates also reduce how quickly those models consume paid ChatGPT Work and Codex subscription credits. Subscription sticker prices and quota budgets did not change.

Fast mode replaces Priority Processing in the API. OpenAI says GPT-5.6 Sol can run up to 2.5 times faster than Standard at twice the token rate, without changing intelligence. Existing requests using service_tier: "priority" continue to work; developers can also use service_tier: "fast".

OpenAI’s developer rate card lists Fast pricing across all three GPT-5.6 tiers. This is a latency purchase, not a quality upgrade. Use it only where faster completion changes revenue, user experience, incident resolution, or developer throughput.

Who benefits—and who loses

High-volume document processing, classification, routine implementation, support automation, and well-specified agent steps gain most from Luna’s cut. It is now materially cheaper than many small and mid-tier proprietary models while retaining GPT-5.6 tool use and multi-step workflow support.

Terra becomes the safer middle route for work that needs more judgment than Luna but cannot justify Sol on every call. OpenAI says Terra reaches GPT-5.5-level intelligence benchmarks at half the price, but production evaluations still matter more than provider benchmarks.

Sol users get no standard-rate cut. Teams that default every request to Sol may now be overpaying for routine steps. Fast mode can also increase spend quickly if enabled globally instead of only on latency-critical paths. Compare alternatives on the Anthropic pricing page and DeepSeek pricing page.

Update cost models immediately, but do not switch model routes without a canary. Run representative tasks through Luna, Terra, and Sol, then track accepted results, input, cache writes, cache reads, output, retries, tool calls, latency, and human correction time.

Split mixed workflows by difficulty. Use Sol to plan or resolve uncertainty, then test whether Terra or Luna can execute the defined steps. Keep stable instructions and tool schemas in cacheable prefixes so repeated agent context benefits from the read discount.

If you used Priority Processing, verify latency and spend after the automatic Fast-mode migration. Pin explicit service tiers in cost-sensitive systems and alert on unexpected Fast traffic.

Our current Cost-per-Task Labs leaderboard already includes all three models. Sol and Terra scored 49/49 in the latest accepted deterministic run; Luna scored 48/49. On July 30, we recalculated Luna and Terra cost-per-correct-answer from that accepted run’s reported token counts using the new direct-provider rates. The outputs and accuracy were not rerun, so the page separates the inference date from the pricing-refresh date.

Bottom line

Luna is the clear economic winner: an 80% rate cut makes it OpenAI’s high-volume GPT-5.6 route. Terra’s 20% reduction strengthens the balanced tier. Sol remains the maximum-capability option, with Fast mode available when latency is worth paying twice the standard rate.

For most API teams, the practical move is routing—not a wholesale migration. Reserve Sol for hard decisions, use Terra for balanced production work, and push defined high-volume steps to Luna after they pass evaluation.

Sources: OpenAI’s price-performance announcement, GPT-5.6 launch, developer pricing documentation, and the live AI Pricing Guru API dataset.