Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.
analysis

GPT-5.6 Sol Finance Benchmark: Model ML Cost Impact

Model ML says GPT-5.6 Sol used fewer tokens for finance decks and workbooks. See the results, current API pricing, caveats, and buyer math.

By AI Pricing Guru Editorial Team

AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.

TL;DR

  • OpenAI reports that Model ML's GPT-5.6 Sol workflow used 21% fewer tokens per PowerPoint deck than Claude Fable 5 and 36% fewer per Excel workbook than Claude Opus 5.
  • Sol produced a PowerPoint file in 100% of tests and cleared Model ML's professional-readiness gate in 43.3%, versus 76% and 26.7% for Opus 5.
  • This is a customer benchmark, not a rate cut: OpenAI's official Sol API price card is unchanged and the live rates below remain canonical.
  • Model ML did not publish the input/output token split, API dollar ledger, prompts, files, or complete grader, so token efficiency cannot be converted into a verified cost-per-deck claim.

GPT-5.6 Sol, Opus 5, and Fable 5 token rates

USD per 1M tokens. Input and output rates are charted separately.

InputOutput
0$50.00GPT 5.6 Solopenai$5.00$30.00Opus 5anthropic$5.00$25.00Fable 5anthropic$10.00$50.00

Estimate your finance-agent token bill

Assumes 75% input tokens and 25% output tokens using current per-million rates.

Claude Opus 5

anthropic

$100.00

Input share
$37.50
Output share
$62.50

GPT-5.6 Sol

openai

$112.50

Input share
$37.50
Output share
$75.00

Claude Fable 5

anthropic

$200.00

Input share
$75.00
Output share
$125.00

Current API rates for the compared models

Model Provider Input / 1M Cached / 1M Output / 1M
GPT-5.6 Sol openai $5.00 $0.5 $30.00
Claude Opus 5 anthropic $5.00 $0.5 $25.00
Claude Fable 5 anthropic $10.00 $1.00 $50.00

Built from pricing.json at publish time.

OpenAI published a Model ML customer story on August 10 showing GPT-5.6 Sol completing native PowerPoint and Excel finance workflows with fewer tokens than selected Claude models. The strongest result is not that Sol has the lowest list price. It is that Model ML’s agent produced more review-ready decks while using fewer tokens than Fable 5, and used substantially fewer tokens than Opus 5 on Excel.

This is new workload evidence, not a new OpenAI model or price change. The live rates above come from our daily pricing dataset, and the official OpenAI price card still matches the canonical Sol row. Buyers should compare the OpenAI pricing page with Anthropic pricing before turning token counts into budget forecasts.

What Model ML measured

Model ML builds agents that research, calculate, and create editable finance deliverables. Its Composite evaluation follows a task from a finance brief through calculations and source reconciliation to a native PowerPoint deck or Excel workbook. The company says its PowerPoint sample spans hundreds of generated decks.

Model ML metricGPT-5.6 SolClaude Opus 5Claude Fable 5
PowerPoint deck produced100.0%76.0%82.0%
PowerPoint professional-readiness rate43.3%26.7%32.0%
PowerPoint readiness-gated quality59.9%56.7%59.3%
Tokens per PowerPoint deck1.10M953K1.40M
Excel headline outputs correct83.3%82.8%80.6%
Excel models with every key output correct50.0%60.0%60.0%
Tokens per Excel workbook2.44M3.83M2.59M
Minutes per Excel workbook7.07.58.4

Sol used about 21% fewer tokens per deck than Fable 5 and about 36% fewer tokens per workbook than Opus 5. It also led Opus 5’s PowerPoint professional-readiness rate by 16.6 percentage points.

The trade-off matters. Opus 5 used fewer tokens per PowerPoint deck than Sol, and both Claude models had a higher share of fully correct Excel models. Sol’s case is strongest where file delivery and review readiness matter alongside correctness—not where a buyer wants a universal accuracy winner.

Pricing impact: token efficiency is not invoice cost

OpenAI announced no rate change with the case study. The official pricing documentation continues to list Sol’s standard input, cached-input read, cache-write, and output rates; the canonical/API Sol row now publishes the cache-write rate and the separate long-context tier as structured fields alongside the core rates shown above.

Model ML reports total tokens per deliverable, but does not disclose how those totals divide among input, cached input, cache writes, and output. Those units have different rates. The story also omits tool charges, web search, document processing, retries, fallback calls, Model ML subscription fees, and the complete API dollar ledger.

That makes an exact “cost per deck” calculation impossible from the published evidence. A buyer can reproduce the math only after recording each billing category from a real trace. Use the calculator above for scenarios, but do not enter 1.10M as though every token were billed at one rate.

The customer story separately says one asset manager reduced a bespoke tearsheet workflow from about one hour to about five minutes. That is a reported operational time saving, not evidence that every finance team will reduce labor cost by the same amount.

What the benchmark does and does not prove

The evaluation is more useful than a generic chat benchmark because it checks editable files, formulas, sources, visual hierarchy, and deliverability. These are real failure points in investment decks and financial models.

However, OpenAI and Model ML did not publish the full prompt set, source documents, generated files, model settings, judge prompts, sample counts for every slice, variance, or a runnable reproduction package. Model ML’s document tooling and agent harness are part of the measured system. The results therefore support a Model ML workflow claim, not a universal ranking of the base models.

The source also gives mixed signals that buyers should preserve. Sol produced every tested deck and led professional readiness, but Opus 5 had lower PowerPoint token use and a higher fully-correct Excel rate. Fable 5 nearly matched Sol’s readiness-gated deck quality. Routing by task and acceptance threshold is more defensible than replacing every finance workflow with one model.

Labs status: Sol included, finance workflow blocked

AI Pricing Guru Labs already includes GPT-5.6 Sol, Claude Opus 5, and Claude Fable 5 in its deterministic 49-task cost-per-correct-answer suite. In the latest accepted public result, Sol answered 49 of 49 tasks correctly with zero API errors; Opus 5 answered 48 and Fable 5 answered 46.

Those results confirm callable routes and provide a separate small-task cost comparison. They do not reproduce Model ML’s finance benchmark. Our suite does not generate native .pptx or .xlsx files, run Model ML’s private agent harness, inspect formulas across workbooks, or apply its professional-readiness grader.

A launch-specific Labs reproduction is explicitly blocked until the evaluation files, fixed prompts, document tools, model settings, grading rubric, and complete per-category usage logs are available. Re-labeling a generic text result as finance-agent evidence would overstate what was tested.

What finance teams should do

  1. Build a blind set of real decks and workbooks with approved answers, formulas, and source citations.
  2. Score file delivery, key-output accuracy, formula integrity, traceability, formatting, and review time separately.
  3. Log input, cached input, cache writes, output, tool calls, retries, and fallbacks for every deliverable.
  4. Calculate cost per accepted file, not cost per initial generation or price per million tokens.
  5. Route by task: one model may be better for presentations while another wins strict spreadsheet correctness.
  6. Keep a finance professional responsible for assumptions, evidence, and final sign-off.

For a broader model-level comparison, use the token cost calculator and the GPT-5.6 pricing analysis.

Bottom line

Model ML’s benchmark makes a credible case that GPT-5.6 Sol can be efficient in a document-heavy finance agent. It produced more decks, cleared professional readiness more often than Opus 5, and used fewer tokens than Fable 5 on PowerPoint and Opus 5 on Excel.

It does not show a new list price or a universally cheaper finance model. The missing billing mix and reproduction package prevent an exact cost-per-deliverable claim. Treat the result as a strong pilot target, then measure accepted-file cost and human review time on your own documents.

Sources: OpenAI, Model ML completes finance work more efficiently with GPT-5.6 Sol, and the official OpenAI API pricing documentation. Published and verified August 10, 2026.