Home Pricing OpenAI

OpenAI API Pricing (August 2026)

Last updated:

How much does the OpenAI API cost? The current public GPT-5.6 ladder runs from Luna at $0.20 input / $1.20 output per 1M tokens to Sol at $4/$20. OpenAI labels Sol's new rate promotional through at least November 21, 2026. Controlled-access GPT-5.6 Cyber and GPT-5.5 Cyber remain at $12.50 input / $75 output.

All OpenAI Models — Price per 1M Tokens

Showing 11 current models. .
Relative priceLowestLowerHigherHighestAutomatic log scale across current offers; input, cache and output are graded separately. Color shows cost, not quality.
Showing 11 grouped models from 11 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

OpenAI remains the largest AI API provider in 2026. Its current published GPT-5.6 ladder runs from GPT-5.6 Luna at $0.20 per million input tokens to GPT-5.6 Sol at $4.00 per million input tokens. Older GPT rows stay in our table as legacy history so migrations and old budgets remain explainable.

The biggest OpenAI pricing story right now is the August 22 GPT-5.6 Sol price cut. Standard input, cached-input, and cache-write rates fell 20%; output fell 33.3%. OpenAI calls the $4/$20 short-context rate promotional and says it will remain available at least through November 21, 2026. Read the full Sol price-cut and workload-cost analysis.

GPT-5.6 is now available in Kiro

OpenAI and AWS added GPT-5.6 Sol, Terra, and Luna to Kiro on August 24. Kiro bills these models in product credits rather than the direct API token rates above: its current model table lists Sol at 2.4× Auto, Terra at 1.0×, and Luna at 0.1×. The models require a paid Kiro plan, and Kiro says GPT-5.6 requests are served from the US even for European profiles.

OpenAI reports that Terra completed successful Terminal-Bench 2.1 tasks in Kiro at roughly 82% lower cost, but the announcement does not identify the comparison baseline or publish a reproducible task ledger. Treat it as a vendor-run cost-per-success claim, not an API price cut or an 82% invoice guarantee. Read the GPT-5.6 in Kiro pricing, credit, and benchmark analysis.

OpenRouter keeps a separate 50% Sol promotion

OpenRouter began advertising a 50% discount on eligible GPT-5.6 Sol routes on August 17. After OpenAI's direct cut, its endpoints API now shows the OpenAI Standard route at $2 input, $0.20 cached input, $2.50 cache write, and $10 output per 1M short-context tokens. Its long-context route is $4/$0.40/$5/$15.

OpenAI direct is now $4/$0.4/$5/$20 for short context and $8/$0.80/$10/$30 above 272K input tokens. Azure, Bedrock, BYOK, regional, Batch, Flex, and Fast-mode invoices can differ, so attach the provider and service tier to every price comparison.

For 100M uncached input plus 20M output tokens, OpenRouter's advertised Standard route is $400 versus $800 at OpenAI direct list price. Read the separate GPT-5.6 Sol OpenRouter discount analysis.

A community oh-my-pi coding stack uses Luna selectively for vision while keeping text and code on DeepSeek V4 Flash. It is a routing pattern, not a new OpenAI price change; the table above remains the official current Luna rate.

GPT-5.6 Family

GPT-5.6 Sol, Terra, and Luna are now generally available through the OpenAI API, ChatGPT, and Codex:

  • GPT-5.6 Sol ($4.00 input / $20.00 output, $0.40 cached input reads) — flagship model for hard reasoning, coding, cybersecurity, and agentic tasks.
  • GPT-5.6 Terra ($2.00 input / $12.00 output, $0.20 cached input reads) — balanced tier for premium production workloads.
  • GPT-5.6 Luna ($0.20 input / $1.20 output, $0.02 cached input reads) — lower-cost tier for faster everyday usage.

Read the current buyer analysis in OpenAI's GPT-5.6 price cut: impact and what it means.

For workload evidence beyond generic coding scores, see the GPT-5.6 versus Claude Fable 5 physical-AI cost test. JuliaHub's sealed simulation study put Sol second on score and first on value among the premium routes it tested.

GPT-5.6 Sol vision benchmark

Roboflow's July vision study resurfaced on Hacker News on August 17. It measured Sol at 46.2 mAP@50 for object detection versus 13.8 for GPT-5.5, and 73.0% counting accuracy versus 64.9%. The same report put GPT-5.5 slightly ahead on OCR and 5.1 percentage points ahead on targeted extraction, so “best vision model” is an overall Roboflow judgment rather than a clean sweep.

Roboflow estimated about 2.5 cents and 10 seconds per image for Sol in its harness. Those are workload observations, not a new OpenAI rate: the maintained Sol row remains the official Standard token price shown above. Image size, reasoning effort, prompt format, token mix, retries, and service tier can change per-image cost.

Labs already includes Sol in the 49-task model-level leaderboard, but it does not reproduce this vision study. A vision replay is explicitly blocked until the promised full test set, fixed prompts, image preprocessing, reasoning settings, raw traces, and bounding-box/OCR graders are public. Read the full GPT-5.6 Sol vision benchmark and pricing analysis.

GPT-5.6 builder guide: lower the cost per accepted task

OpenAI's August 13 builder guide does not change the GPT-5.6 rate card. It recommends testing lower reasoning effort, routing routine steps to Luna or Terra, preserving reasoning across Responses API calls, compacting long histories, moving deterministic filtering into programmatic tool code, and using multi-agent execution only where parallel work earns back the extra token spend.

The guide says GPT-5.6 Luna scored 84.04% on BrowseComp at a reported $1.33 benchmark cost, close to GPT-5.5's 84.36% at $33.27. It also reports that retained reasoning and compaction moved Sol from 13.3% to 38.3% on ARC-AGI-3 while using roughly six times fewer output tokens. These are OpenAI-run examples, not universal savings guarantees.

Prompt-cache time to live is now at least 30 minutes across the family, with deterministic cache breakpoints available. Buyers should validate cache-hit rate, retries, latency, human correction, and total cost on their own accepted task set. Read the full GPT-5.6 builder guide pricing analysis.

GPT-5.6 Sol Ultrafast preview

OpenAI's August 13 Ultrafast preview runs GPT-5.6 Sol on Cerebras infrastructure at up to 750 output tokens per second and up to 14× Standard processing speed. Access is limited to a select customer group while capacity expands.

No public Ultrafast rate, quota, region list, minimum commitment, or service-level term has been published. The Sol row above therefore remains the verified Standard rate—not an Ultrafast estimate. Fast mode is a separate generally documented tier at 2× the Standard token rate and up to 2.5× Standard speed.

Raw output throughput is not end-to-end task latency. Buyers should compare time to first token, tool and network time, retries, human correction, and cost per accepted task on the same workload before paying for a premium tier. Read the full GPT-5.6 Sol Ultrafast pricing analysis.

GPT-5.6 Cyber and Daybreak pricing

OpenAI's cyber rate card lists GPT-5.6 Cyber at $12.5 input, $1.25 cached input, $15.625 cache write, and $75 output per 1M tokens. Against Sol's new promotional rate, Cyber is 3.125x the input-side categories and 3.75x output.

Daybreak Blue gives approved defenders GPT-5.6 Sol for vulnerability discovery, secure code review, malware analysis, incident response, and patch validation. Daybreak Red provides purpose-trained GPT-5.6 Cyber for authorized vulnerability research, exploit validation, and security testing.

OpenAI's internal Advanced Cybersecurity Completion Rate reports 95.0% for GPT-5.6 Cyber through Red, 57.3% for GPT-5.5 Cyber through Red, 2.0% for Sol through Blue, and 1.5% for standard Sol. This measures whether the model responds to advanced requests—not answer accuracy, exploit validity, safety, or successful remediation.

The token rate is a model-cost baseline, not a self-serve promise or a full customer quote. OpenAI limits Daybreak access to approved defenders and partners, while products and engagements can add platform, governance, and expert-service fees. Read the GPT-5.6 Cyber and Daybreak Blue versus Red analysis for pricing, buyer guidance, and the Labs access blocker.

Model ML finance benchmark

OpenAI's August 10 Model ML customer story adds workload-specific evidence for native PowerPoint and Excel creation. In Model ML's agent harness, GPT-5.6 Sol used 1.10M tokens per deck—21% fewer than Claude Fable 5—and 2.44M tokens per workbook—36% fewer than Claude Opus 5. Sol produced a deck in 100% of tests and cleared the professional-readiness gate in 43.3%, versus 76% and 26.7% for Opus 5.

This did not change the official rate card: Sol remains $4 input / $0.4 cached input / $5 cache write / $20 output per 1M short-context tokens. Above 272,000 input tokens, the maintained row publishes $8 / $0.8 / $10 / $30. Model ML did not publish the billing split, full dollar ledger, or reproduction package, so token totals are not an exact cost-per-deck result. Read the Model ML finance benchmark cost analysis for the comparison table, Labs boundary, and buyer checklist.

GPT-Realtime-2.1 Pricing

GPT-Realtime-2.1 is OpenAI's current speech-to-speech reasoning model for the Realtime API. It accepts text, audio, and images; produces text and audio; supports tool use and prompt caching; and has a 128K context window with up to 32K output tokens.

Modality Input / 1M Cached input / 1M Output / 1M
Audio $32.00 $0.40 $64.00
Text $4.00 $0.40 $24.00
Image $5.00 $0.50 n/a

Realtime billing accumulates across conversation turns, so later responses can include prior text, audio, and image context. OpenAI says user audio represents one token per 100 milliseconds and assistant audio one token per 50 milliseconds. Measure response.done usage rather than treating these rates as a flat per-minute fee.

OpenAI's Avatarin customer story shows the model in a 24/7 multilingual retail agent used by roughly 30,000 people in two weeks. Read our GPT-Realtime retail-agent cost analysis for deployment economics and the missing campaign-spend caveats.

ARC-AGI-3: Retained Reasoning and Compaction

OpenAI's ARC-AGI-3 study shows why agent harness design affects both benchmark scores and token bills. On the same public set, GPT-5.6 Sol moved from 13.3% RHAE with the generic harness to 38.3% when the Responses API retained reasoning across turns and compacted long histories. OpenAI also reports 6x fewer output tokens. Read our ARC-AGI-3 cost analysis for the billing caveats and reproducibility limits.

GPT-5.6 Context and Long-Context Pricing

All three GPT-5.6 tiers have a 1.05M-token context window, a 922K maximum input, and a 128K maximum output. Requests with more than 272K input tokens are billed at 2x the standard input rate and 1.5x the standard output rate for the full request. The ARC-AGI-3 Responses harness used a 175K-token compaction threshold, below that long-context price boundary.

Legacy GPT-5.5 Family

GPT-5.5 is no longer present in the latest official OpenAI pricing acquisition, so we keep these rows as legacy history:

  • GPT-5.5 ($5.00 input / $30.00 output, $0.50 cached input) — new flagship for complex professional work, coding, and long-context agents.
  • GPT-5.5 Pro ($30.00 input / $180.00 output) — highest-precision variant for expensive-but-important workloads. No cached-input discount.

Legacy GPT-5.4 Family

GPT-5.4 is also retained as legacy history after the current OpenAI pricing page moved to GPT-5.6:

  • GPT-5.4 ($2.50 input / $15.00 output) — superseded in defaults by GPT-5.6 Terra.
  • GPT-5.4 mini ($0.75 input / $4.50 output) — retained for old budget comparisons.
  • GPT-5.4 nano ($0.20 input / $1.25 output) — retained for old routing and extraction comparisons.

Legacy o-Series Reasoning Models

OpenAI's older reasoning rows remain useful for historical cost comparisons, but they are hidden behind the legacy toggle when they are no longer part of the current default price card.

Legacy GPT-4.1 Family

The GPT-4.1 family is now treated as legacy in AI Pricing Guru defaults because it no longer appears in the latest OpenAI pricing acquisition:

  • GPT-4.1 ($2.00 input / $8.00 output) — 1M context, strong for long-document processing
  • GPT-4.1 mini ($0.40 input / $1.60 output) — Best value for large context needs
  • GPT-4.1 nano ($0.10 input / $0.40 output) — Cheapest model in OpenAI's lineup

Cached Input Pricing

GPT-5.6 supports cached-input reads at 90% off the standard input rate. Cache writes for GPT-5.6 and later models cost 1.25x the uncached input rate: $5/M for Sol, $2.50/M for Terra, and $0.25/M for Luna. Sol's cached read rate is $0.4/M, Terra's is $0.20/M, and Luna's is $0.02/M.

How OpenAI Compares

OpenAI now covers several premium price bands. GPT-5.6 Sol's promotional $4/$20 rate is below Claude Opus 4.8's $5/$25, while GPT-5.6 Luna is cheaper than Claude Sonnet 5's current $2/$10 rate. DeepSeek still undercuts both on pure token price.

For budget use cases, compare GPT-5.6 Luna with Google's Gemini Flash tiers, DeepSeek, and hosted open models before choosing a default route.

Price History

Only models with a recorded price change are charted here.

GPT-5.6 Luna

GPT-5.6 Sol

GPT-5.6 Terra

Price history tracking started April 2026. Flat model charts stay hidden until a price change is detected.
View pricing changelog →

Frequently asked questions

How much does GPT-5.6 cost?

GPT-5.6 has three tiers: Sol costs $4.00 per 1M input tokens and $20.00 per 1M output tokens, Terra costs $2.00/$12, and Luna costs $0.20/$1.20. Sol's promotional pricing is available at least through November 21, 2026.

How much does GPT-5.6 Cyber cost?

OpenAI lists GPT-5.6 Cyber at $12.50 input, $1.25 cached input, $15.625 cache write, and $75 output per 1M tokens. Access is controlled through approved Daybreak partners rather than a public self-serve route.

What is the difference between Daybreak Blue and Red?

Daybreak Blue gives approved defenders GPT-5.6 Sol with safeguards tailored to authorized defensive work. Daybreak Red provides purpose-trained cyber models such as GPT-5.6 Cyber for closely governed vulnerability research, exploit validation, and security testing.

How much does GPT-5.5 cost per token?

The last retained GPT-5.5 price in our history is $5.00 per 1M input tokens and $30.00 per 1M output tokens. OpenAI no longer publishes GPT-5.5 on the current API pricing page, so it is shown as legacy in the table.

Does OpenAI have a free tier?

OpenAI does not offer an ongoing free API tier for production use. New accounts typically receive starter credits, but serious usage is paid. For free experimentation, Google Gemini still offers limited Flash and Flash-Lite access — see our Google AI pricing page.

How does OpenAI compare to Anthropic?

At the flagship tier, GPT-5.6 Sol now costs $4/$20, while Claude Opus 4.8 is $5/$25. GPT-5.6 Terra costs $2/$12, close to Claude Sonnet 5 at $2/$10.

What models does OpenAI offer?

OpenAI now publishes GPT-5.6 Sol/Terra/Luna on its current API pricing page. Older GPT-5.5, GPT-5.4, GPT-4.1, GPT-4o, and o-series rows remain in AI Pricing Guru as legacy history and compatibility references.

Is GPT-5.6 generally available?

Yes. OpenAI made GPT-5.6 Sol, Terra, and Luna generally available through the API, ChatGPT, and Codex on July 29, 2026. Account-level rate limits and product rollout timing can still vary.

What's GPT-5.6's context window?

GPT-5.6 Sol, Terra, and Luna each have a 1.05M-token context window, with up to 922K input tokens and 128K output tokens. OpenAI charges 2x input and 1.5x output for the full request when input exceeds 272K tokens.

How cheap is GPT-5.6 Luna?

GPT-5.6 Luna is the cheapest current GPT-5.6 tier at $0.20 per 1M input tokens and $1.20 per 1M output tokens after OpenAI's July 30 price cut.

How much does GPT-5.6 Sol Ultrafast cost?

OpenAI has not published an Ultrafast token rate. The limited preview runs GPT-5.6 Sol at up to 750 output tokens per second and up to 14x Standard speed, but the maintained Sol row still represents Standard pricing. Fast mode is separately published at 2x the Standard token rate.

How much does GPT-Realtime-2.1 cost?

GPT-Realtime-2.1 costs $32 per 1M audio input tokens, $0.40 cached, and $64 per 1M audio output tokens. Text costs $4 input, $0.40 cached, and $24 output; image input is $5, or $0.50 cached.

Does Model ML prove GPT-5.6 Sol is the cheapest model for finance?

No. Model ML reports fewer tokens for selected PowerPoint and Excel comparisons and stronger deck delivery, but it did not publish the billing-category split or a complete dollar ledger. The result supports a workflow pilot, not a universal cheapest-model claim.

Is GPT-5.6 Sol OpenAI's best vision model?

Roboflow's July benchmark calls Sol OpenAI's strongest vision model overall, with a large detection and counting gain over GPT-5.5. It is an independent workload result, not an OpenAI claim: GPT-5.5 still scored slightly higher on OCR and materially higher on targeted extraction, while Gemini 3.5 Flash led Roboflow's detection-and-cost trade-off.

Did GPT-5.6 Sol get a price cut?

Yes. On August 22, OpenAI cut direct Standard input pricing 20% to $4 and output pricing 33.3% to $20 per 1M short-context tokens. OpenRouter separately maintains a 50% promotion on eligible OpenAI routes, currently $2/$10. Route, cloud, regional, Batch, Flex, Fast mode, and BYOK pricing can differ.

Methodology

Pricing sourced from OpenAI's developer pricing documentation on . All token prices are USD per 1 million tokens. Raw data: /api/pricing.json. API docs.

Compare All Providers

See how OpenAI stacks up against Anthropic, Google, DeepSeek, and more.

Further reading: OpenAI vs Anthropic pricing · ChatGPT vs Claude · Best AI API for developers. Looking for voice/TTS API pricing? Compare ElevenLabs, Speechify, OpenAI audio, Google TTS, and Amazon Polly.