GPT-5.6 vs Claude Fable 5: Physical AI Cost Test
Fable 5 led JuliaHub's physical-AI test, but GPT-5.6 Sol and Terra cost far less per trial. Compare scores, API prices, limits, and buyer fit.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- Claude Fable 5 won JuliaHub's sealed physical-AI evaluation with a 0.889 weighted score, ahead of GPT-5.6 Sol at 0.814.
- The quality lead was expensive: Fable averaged $9.60 per trial versus $1.74 for Sol and $1.25 for Terra.
- For maximum engineering accuracy, Fable won this test. For measured score per dollar, GPT-5.6 Terra and Sol were the stronger starting points.
- This is one vendor's 52-run simulation study, not a universal model ranking; the hardest flight problem defeated every model.
API workload cost at current list prices
USD per 1M tokens. Input and output rates are charted separately.
Calculate your physical-AI agent token bill
Assumes 75% input tokens and 25% output tokens using current per-million rates.
GPT-5.6 Luna
openai
$4.50
- Input share
- $1.50
- Output share
- $3.00
GPT-5.6 Terra
openai
$45.00
- Input share
- $15.00
- Output share
- $30.00
GPT-5.6 Sol
openai
$80.00
- Input share
- $30.00
- Output share
- $50.00
Claude Fable 5
anthropic
$200.00
- Input share
- $75.00
- Output share
- $125.00
Official Claude Fable 5 and GPT-5.6 API prices
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| Claude Fable 5 | anthropic | $10.00 | $1.00 | $50.00 |
| GPT-5.6 Sol | openai | $4.00 | $0.4 | $20.00 |
| GPT-5.6 Terra | openai | $2.00 | $0.2 | $12.00 |
| GPT-5.6 Luna | openai | $0.2 | $0.02 | $1.20 |
Built from pricing.json at publish time.
Editor’s note (updated July 31, 2026): JuliaHub’s trials ran before OpenAI’s July 30 GPT-5.6 price cut, which reduced Luna standard rates by 80% and Terra by 20% (Sol unchanged). The list prices quoted in “Why Fable cost more” reflect the pre-cut card in force at the time of the test; the live table, chart, and calculator above render today’s rates from pricing.json, so trial-cost gaps against Terra and Luna would be narrower today than JuliaHub reported. See our GPT-5.6 price cut analysis for the new rate ladder.
Claude Fable 5 performed best in JuliaHub’s new Physical AI evaluation, but GPT-5.6 Sol and Terra delivered much lower trial costs. The result gives engineering teams a useful routing signal: pay for Fable when a modest quality gain changes the outcome, and start with Sol or Terra when many simulations must run inside a fixed budget.
JuliaHub tested four models inside the same Dyad agent harness on five modeling and simulation problems. The harness, sealed tasks, one-million-token context setting, 128K token budget, and xhigh reasoning configuration stayed fixed. The study ran 52 graded trials in total.
Physical AI benchmark results
| Model | Weighted score | Average cost per trial | Average time |
|---|---|---|---|
| Claude Fable 5 | 0.889 | $9.60 | 16.1 min |
| GPT-5.6 Sol | 0.814 | $1.74 | 13.4 min |
| GPT-5.6 Terra | 0.786 | $1.25 | 12.6 min |
| GPT-5.6 Luna | 0.727 | $3.26 | 25.0 min |
Fable’s weighted score was about 9% higher than Sol’s, while its average trial cost was roughly 5.5 times higher. Terra cost the least and finished fastest. Luna trailed the other three on score, cost, and time in this specific setup.
That makes the verdict workload-dependent:
- Best measured physical-AI quality: Claude Fable 5.
- Best balance of score and cost: GPT-5.6 Sol.
- Lowest measured trial cost and time: GPT-5.6 Terra.
- Weakest result in this harness: GPT-5.6 Luna.
Do not convert the weighted score directly into a universal “accuracy” percentage. JuliaHub combines five task scores with more weight on the hardest problems. Cost and time are separate unweighted means.
What the models actually had to do
The tasks were not simple physics questions. The agent had to read engineering specifications, derive equations, write and compile a Dyad model, simulate it, verify the output, and compare the trajectory with sealed ground truth.
The hardest problem modeled NASA’s HL-20 lifting-body vehicle with six-degree-of-freedom dynamics, aerodynamic tables, atmosphere behavior, and control surfaces. No model solved it. Fable scored 0.69 on that problem, Sol 0.66, Terra 0.59, and Luna 0.38.
That limitation matters. The study supports a narrower claim—Fable was strongest among these four configurations—not a claim that today’s frontier models can autonomously replace simulation engineers.
Why Fable cost more
The official list price explains much of the gap. Anthropic charges $10 per million input tokens and $50 per million output tokens for Fable 5. OpenAI lists GPT-5.6 Sol at $5/$30, Terra at $2.50/$15, and Luna at $1/$6.
Token rates alone do not reproduce JuliaHub’s trial bills. The models used different numbers of tokens and tool calls, and agent behavior affected runtime. Luna is the clearest warning: it has the cheapest GPT-5.6 list price, yet it averaged the highest GPT trial cost because it iterated more and took longer.
For physical-AI agents, track:
- successful simulated trajectories, not only compilations;
- input, reasoning, and output tokens per accepted result;
- retries and tool calls;
- wall-clock time and compute occupied by external simulators;
- expert review needed before a model is trusted.
The lowest price per million tokens can lose when it produces more failed loops.
How this compares with AI Pricing Guru Labs
All four models are already included in the current AI Pricing Guru Labs leaderboard, but our suite answers a different question. Labs uses 49 deterministic short tasks covering exact answers, extraction, structured output, and pricing logic. It does not run JuliaHub’s private Dyad harness or sealed physical-simulation tasks.
In the July 29 Labs run, GPT-5.6 Sol and Terra each scored 49/49, Luna scored 48/49, and Fable scored 46/49, all with zero API errors. Those results do not contradict JuliaHub. They show that model rankings can reverse when the task class changes.
The explicit availability blocker is the Physical AI experiment itself: JuliaHub’s sealed problems, trajectory grader, Dyad configuration, and transcripts are not part of our independent Labs adapter. We therefore report its published methodology and results without presenting them as reproduced.
Which model should engineering teams choose?
Choose Fable 5 when a wrong physical model creates a large downstream cost and the task resembles long-horizon engineering work. Its extra $7.86 per average trial over Sol is small if it prevents hours of expert correction or a flawed design decision.
Choose GPT-5.6 Sol when you want most of the measured quality at much lower cost. It came second overall, finished faster than Fable, and stayed close on the hardest flight task.
Choose Terra for high-volume exploration, early-stage derivations, or parallel candidate generation. Escalate uncertain or high-value cases to Sol or Fable. That tiered router is likely more economical than sending every problem to the most expensive model.
Treat Luna cautiously for this workload. Its low official token rate did not translate into low agent cost in JuliaHub’s configuration.
If you want a lower-cost open-model control group before operating GPUs yourself, benchmark a managed route on Novita against the same sealed tasks. We may earn a commission from that sponsored link; it does not affect the benchmark interpretation.
Bottom line
Claude Fable 5 won this Physical AI test on quality. GPT-5.6 Sol won the practical value argument, and Terra was the cheapest and fastest trial route. The strongest production design is not one universal default: use cheaper models to explore and reserve Fable for problems where its measured quality edge survives your own sealed evaluation.
Compare the full rate cards on our OpenAI pricing page and Anthropic pricing page, then model your token mix in the AI cost calculator.
Sources: JuliaHub’s Physical AI evaluation, OpenAI’s GPT-5.6 announcement and pricing, Anthropic’s Claude Fable 5 page, and the live AI Pricing Guru API dataset.