OpenAI Third-Party Assessments — Pricing Impact (Sep 2026)
OpenAI proposes four priorities and seven principles for third-party AI safety assessments. See the assurance cost, limits, and buyer checklist.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- OpenAI proposes independent assessment of safety cases, safeguards, frontier-risk evaluations, and critical misalignment incidents.
- Its seven principles call for pre-registered claims, proportionate access, transparent methods, expert independence, strong security, actionable remediation, and responsible publication.
- No model, subscription, or API price changed. The cost impact is a new assurance layer: specialist assessors, secure environments, staff access, remediation, and recurring reviews.
- Buyers should not accept 'third-party assessed' as a standalone claim—ask what was tested, what access was granted, what was excluded, and what findings were redacted.
Frontier model rates before independent-assurance costs
USD per 1M tokens. Input and output rates are charted separately.
Estimate model usage inside an evaluation program
Assumes 75% input tokens and 25% output tokens using current per-million rates.
DeepSeek V4.1 Flash
deepseek
$2.63
- Input share
- $1.13
- Output share
- $1.50
GPT-5.6 Sol
openai
$80.00
- Input share
- $30.00
- Output share
- $50.00
GPT-6 Astra
openai
$200.00
- Input share
- $75.00
- Output share
- $125.00
Claude Fable 5.1
anthropic
$200.00
- Input share
- $75.00
- Output share
- $125.00
Current frontier API rates—assessment costs are additional
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| GPT-6 Astra | openai | $10.00 | $1.00 | $50.00 |
| GPT-5.6 Sol | openai | $4.00 | $0.4 | $20.00 |
| Claude Fable 5.1 | anthropic | $10.00 | $0.25 | $50.00 |
| DeepSeek V4.1 Flash | deepseek | $0.15 | $0.0030 | $0.6 |
Built from pricing.json at publish time.
OpenAI has proposed a framework for deeper independent assessment of frontier AI safety claims. It calls for third parties to examine evidence across training, evaluation, internal deployment, external deployment, and serious model-misalignment incidents.
This is not a certification launch or proof that OpenAI’s systems have passed a new audit. It is OpenAI’s statement of priorities and operating principles for future assessments, with multiple third-party proposals now under discussion.
The four areas OpenAI wants assessed
| Priority | Core question | Evidence buyers should expect |
|---|---|---|
| Safety cases | Do the claims and evidence support safe training and deployment? | Assumptions, conditions, gaps, and evidence across the model lifecycle |
| Critical safeguards | Do controls survive realistic adversarial testing? | Grey-box tests, monitoring coverage, containment results, and failure modes |
| Capability evaluations | Do tests cover frontier cyber, chemical, biological, self-improvement, and alignment risks? | Threshold definitions, refreshed evals, saturation checks, and missing behaviors |
| Misalignment incidents | What happened, why, and will remediation prevent recurrence? | Forensics, model-behavior analysis, contributing factors, and remediation testing |
OpenAI says assessments may run for weeks or months and are generally intended to be long-term and launch-agnostic. Pre-deployment work can still inform a release decision, but the proposal is broader than a last-minute model scorecard.
Seven principles—and the accountability tension
OpenAI’s principles are concrete: agree and pre-register the claims, provide proportionate access, publish transparent methods, use qualified and independent assessors, protect sensitive information, produce actionable findings with remediation time, and publish responsibly.
The access commitment is notable. OpenAI says past assessors have received information about technical safeguards, visible chain-of-thought access, confidential data, and internal deployment access for incident-response and monitor red-teaming work.
The tension is that scope, access, remediation, confidentiality, and redactions are negotiated with the lab being assessed. OpenAI says assessors should retain editorial independence and may note substantive redactions, but this remains a vendor-authored proposal—not an independent standard or a completed audit. A strong final report must state what was excluded and how those exclusions limit the conclusion.
What do third-party AI assessments cost?
OpenAI announced no new API rate, subscription fee, assessor price, customer surcharge, or public price for the assessment work itself. The live chart, calculator, and table above show current model rates from AI Pricing Guru’s maintained dataset; they do not include the external-assurance program OpenAI describes.
The additional budget can include:
| Cost layer | Likely work |
|---|---|
| Specialist assessment | Alignment, cyber, bio/chemical, red-team, and forensic expertise |
| Secure access | Controlled devices, environments, logging, confidentiality, and legal review |
| Lab support | Staff time to prepare evidence, answer questions, and enable access |
| Remediation | Engineering fixes, retesting, and proof that controls now work |
| Publication and oversight | Redaction review, correction processes, boards, and regulators |
For enterprise buyers, the practical price is no longer only model tokens. It is tokens plus the cost of establishing that a model and its safeguards are suitable for the intended risk level. Use the token cost calculator for inference, then budget assurance separately. Compare current OpenAI pricing, Anthropic pricing, and DeepSeek pricing without assuming any provider’s public rate includes equivalent third-party scrutiny.
Who benefits—and who bears the cost
Regulated buyers, boards, insurers, and governments benefit if assessments turn broad safety language into claims backed by inspectable evidence. Independent assessors gain a clearer market for deep technical work, while model developers gain a route to challenge internal assumptions before failures become incidents.
Smaller labs may face the highest burden because expert teams and secure access have fixed costs. The public can also lose if confidentiality and redaction make a report sound comprehensive while hiding material exclusions. OpenAI’s principle that no single assessor should cover every frontier question is therefore important: expertise must be distributed, but conclusions must still fit together.
For a non-sensitive prototype evaluation harness, teams can compare DigitalOcean development infrastructure. General cloud hosting is not a substitute for the controlled access, confidentiality, or safety boundaries required for frontier-model assessment.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored link above at no extra cost to you. It does not affect this analysis.
What procurement teams should require
- List every safety claim, model version, deployment surface, and excluded system.
- Record assessor funding, prior lab relationships, conflicts, recusals, and independence safeguards.
- Explain direct versus indirect access and any lab-controlled evidence collection.
- Publish methods, criteria, uncertainty, negative results, and the boundary between findings and interpretation.
- Set remediation and retest dates, then define what triggers another assessment after a model or safeguard change.
- Disclose requested and accepted redactions and explain how they affect the conclusion.
Our earlier SWE-Bench Pro audit analysis shows why methodology matters even for narrower performance tests. This policy announcement does not trigger an AI Pricing Guru Labs run because it makes no new model-performance claim, model artifact, or callable route that a prompt benchmark can verify. A defensible assessment study would instead need a pre-registered safety claim, the exact model and safeguard version, authorized grey-box access, fixed adversarial cases, secure evidence handling, complete traces, and an independent publication process.
Bottom line
OpenAI’s proposal is a useful step toward making frontier safety claims testable by outsiders. Its strongest elements are pre-registered claims, proportionate access, conflict disclosure, explicit uncertainty, remediation, and visible redaction boundaries.
No API price changed. The economic change is that credible frontier deployment increasingly requires an assurance budget alongside inference spend—and buyers need the report scope before they can decide whether that assurance is real.
Sources: OpenAI’s priorities and principles for third-party assessments, Preparedness Framework update, and AI Pricing Guru’s live pricing dataset. Proposal and rates checked September 22, 2026 at 17:48 UTC.