AI Pricing Guru Labs · Experiment 01

Which AI model gives the most correct answers per dollar?

We ran 15 models through the same 49 short, machine-graded tasks. The primary score is cost per correct answer. A model must score at least 95% to qualify, so a low price cannot compensate for too many wrong answers.

See the verdict Compare all models Open prompts + raw results

Live experiment 01

One question, one primary metric

Question: for routine, short-form API work, which model produces correct answers at the lowest token cost? We divide each model’s total direct-price run cost by the number of correct answers. Lower is better—but only if the observed error rate is acceptable for your workflow.

Primary score: cost per correct answer. This combines the bill and the observed success rate into one number. Qualified models are sorted by this score, so there are no confusing accuracy ties.
Quality guardrail: at least 95% accuracy. A model needs 47/49 correct to receive a value rank. Lower-scoring models remain visible, but they are marked below threshold rather than recommended.
15 models in this run Task suite v2 Full benchmark 49 tasks per model 95% ranking threshold 735 answers graded Run 2026-07-29 Costs repriced 2026-07-30

The verdict

What works, what does not, and what it costs

Best value for this workload

DeepSeek V4 Flash (pre-0731)

48/49 correct · 98% accuracy

$0.001526 for all 49 tasks · $0.000032 per correct answer. Perfect in Reasoning, Extraction, Summarization, Coding, Instruction, Pricing. Missed: Classification: 1 missed.

Cheapest zero-miss option

GPT-5.6 Terra

49/49 correct · 100% accuracy

$0.020772 for the run · $0.000424 per correct answer. It bought 1 more correct answers than DeepSeek V4 Flash (pre-0731), but at 13.3× the cost per correct answer.

Poor value on these simple tasks

Gemini 3.6 Flash

49/49 correct · 100% accuracy

$0.099127 for the run · $0.002023 per correct answer, or 63.6× the best-value score. That does not make it a bad model; it means its premium was not justified by this workload.

Decision in plain English

Choose DeepSeek V4 Flash (pre-0731) when occasional misses can be validated or retried. Choose GPT-5.6 Terra when this exact task mix requires zero observed misses. At the same prompt lengths and mix, one million correct answers would cost roughly $32 versus $424. Do not extrapolate that estimate to longer or more complex workloads.

Same-scale comparison

What one million correct answers would cost

Every bar uses the same linear $0–$2.4k axis, so the price gap is proportional rather than decorative. This projects each qualified model’s measured cost per correct answer to one million answers; lower is better.

DeepSeek V4 Flash (pre-0731)

98% accuracy

Projected cost $32 best value

GPT-5.6 Luna

98% accuracy

Projected cost $53 1.7× vs. best

DeepSeek V4 Pro

98% accuracy

Projected cost $129 4.1× vs. best

GPT-5.6 Terra

100% accuracy

Projected cost $424 13.3× vs. best

Claude Sonnet 5

98% accuracy

Projected cost $435 13.7× vs. best

GLM-5.2

98% accuracy

Projected cost $601 18.9× vs. best

GPT-5.6 Sol

100% accuracy

Projected cost $852 26.8× vs. best

Claude Opus 4.8

100% accuracy

Projected cost $859 27× vs. best

Grok 4.5

100% accuracy

Projected cost $1,397 43.9× vs. best

Claude Opus 5

98% accuracy

Projected cost $1,690 53.2× vs. best

Gemini 3.6 Flash

100% accuracy

Projected cost $2,023 63.6× vs. best

This is a straight-line scale-up of the observed 49-task mix, not a volume quote. It excludes models below the 95% quality threshold and should not be extrapolated to longer prompts or different workloads.

Cost-per-task leaderboard

Lowest cost per correct answer wins

Value rank #1

DeepSeek V4 Flash (pre-0731)

$0.000032

per correct

Accuracy98%
Run cost$0.001526
Vs. bestBest

Perfect: Reasoning, Extraction, Summarization, Coding, Instruction, Pricing
Missed: Classification: 1 missed

Value rank #2

GPT-5.6 Luna

$0.000053

per correct

Accuracy98%
Run cost$0.002554
Vs. best1.7×

Perfect: Reasoning, Extraction, Classification, Summarization, Instruction, Pricing
Missed: Coding: 1 missed

Value rank #3

DeepSeek V4 Pro

$0.000129

per correct

Accuracy98%
Run cost$0.006184
Vs. best4.1×

Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Pricing
Missed: Instruction: 1 missed

Value rank #4

GPT-5.6 Terra

$0.000424

per correct

Accuracy100%
Run cost$0.020772
Vs. best13.3×

Perfect across all 7 categories. No misses in this run.

Value rank #5

Claude Sonnet 5

$0.000435

per correct

Accuracy98%
Run cost$0.020864
Vs. best13.7×

Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Instruction
Missed: Pricing: 1 missed

Value rank #6

GLM-5.2

$0.000601

per correct

Accuracy98%
Run cost$0.028862
Vs. best18.9×

Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Instruction
Missed: Pricing: 1 missed

Value rank #7

GPT-5.6 Sol

$0.000852

per correct

Accuracy100%
Run cost$0.04173
Vs. best26.8×

Perfect across all 7 categories. No misses in this run.

Value rank #8

Claude Opus 4.8

$0.000859

per correct

Accuracy100%
Run cost$0.042085
Vs. best27×

Perfect across all 7 categories. No misses in this run.

Value rank #9

Grok 4.5

$0.001397

per correct

Accuracy100%
Run cost$0.06845
Vs. best43.9×

Perfect across all 7 categories. No misses in this run.

Value rank #10

Claude Opus 5

$0.00169

per correct

Accuracy98%
Run cost$0.08111
Vs. best53.2×

Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Instruction
Missed: Pricing: 1 missed

Value rank #11

Gemini 3.6 Flash

$0.002023

per correct

Accuracy100%
Run cost$0.099127
Vs. best63.6×

Perfect across all 7 categories. No misses in this run.

Below 95% threshold

Mistral Large 3

$0.000054

per correct

Accuracy90%
Run cost$0.00238
Vs. best1.7×

Perfect: Extraction, Classification, Pricing
Missed: Reasoning: 2 missed · Summarization: 1 missed · Coding: 1 missed · Instruction: 1 missed

Below 95% threshold

Gemini 3.5 Flash-Lite

$0.000059

per correct

Accuracy92%
Run cost$0.002641
Vs. best1.8×

Perfect: Extraction, Classification, Pricing
Missed: Reasoning: 1 missed · Summarization: 1 missed · Coding: 1 missed · Instruction: 1 missed

Below 95% threshold

Claude Haiku 4.5

$0.000145

per correct

Accuracy94%
Run cost$0.006656
Vs. best4.6×

Perfect: Extraction, Classification, Summarization, Coding, Pricing
Missed: Reasoning: 2 missed · Instruction: 1 missed

Below 95% threshold

Claude Fable 5

$0.002709

per correct

Accuracy94%
Run cost$0.12462
Vs. best85.2×

Perfect: Extraction, Classification, Summarization, Instruction
Missed: Reasoning: 1 missed · Coding: 1 missed · Pricing: 1 missed

Models need at least 95% accuracy (47/49) to qualify. Qualified models are ranked by run cost ÷ correct answers. “Vs. best” shows how many times more each correct answer cost than the winner. Models below the quality threshold remain visible without receiving a value rank.

Methodology

Reproducible, deterministic grading

Same prompts and settings

Every model receives the same system prompt and user tasks through OpenRouter with temperature 0, top_p 1, and a 1024-token output ceiling so reasoning tokens do not truncate the final answer.

No LLM judge

Answers are checked with exact, numeric, regex, contains, and JSON-field graders. This avoids paying another model to make subjective scoring decisions.

Direct-provider price basis

OpenRouter is the inference route. Cost is calculated from the direct-provider list prices in AI Pricing Guru’s daily pricing dataset, using reported prompt and completion tokens. GPT-5.6 Luna and Terra costs were recalculated from the accepted run’s token counts on 2026-07-30; outputs and accuracy were not rerun.

What this does not prove

This compact suite does not measure long-context reliability, safety, creative writing, production latency, tool use, or every coding workload. Results are a workload signal, not a universal model ranking.

Availability blocker · GPT-5.6 Cyber

Daybreak Red is controlled access, not a public benchmark route

GPT-5.6 Cyber is not ranked here because OpenAI supplies it only through the governed Daybreak Red tier for approved defenders. Labs has no authorized, stable endpoint or fixed snapshot for the 49-task suite, and substituting public GPT-5.6 Sol would test a different model. A valid entry requires approved access, a safe cyber-specific task set, fixed settings, and complete usage records. See the Daybreak Blue versus Red pricing analysis →

Availability blocker · GPT-Realtime-2.1

Realtime voice needs a different benchmark harness

GPT-Realtime-2.1 is not ranked here because OpenAI exposes it through v1/realtime, not the text-completion route used by this deterministic suite. A valid test needs fixed voices and codecs, controlled turns and interruptions, speech-quality grading, latency measurement, retail catalog fixtures, and cost per resolved conversation. Substituting a text model would not measure the speech-to-speech product. See the Avatarin retail deployment cost analysis →

Availability blocker · Meta's next open-source model

There is no model artifact to benchmark yet

Meta announced that some open-source model releases will resume soon, but published no model name, weights, callable endpoint, license, context window, or reproducible evaluation package. Labs cannot add an unnamed model or substitute an existing Llama checkpoint. We will evaluate the release after Meta publishes an exact artifact and a supported inference route. See the Meta release-status analysis →

Evaluation blocker · Claude content marking

Anthropic has not published a detector or reproducible support matrix

Existing Claude model cost results remain valid because the marking policy changed no public token rate. Labs cannot yet measure watermark survival, false positives, false negatives, or C2PA preservation: Anthropic has not published the official detector, thresholds, model-by-model support status, or fixed text and file test procedure. A valid study needs those artifacts plus controlled copy, edit, translation, screenshot, re-save, and cloud-route transformations. Generic AI detectors are not a substitute. See the Claude marking buyer-impact analysis →

Scope blocker · Claude Code plan economics

Labs can price the models, but not reproduce a private seat allowance

Claude Sonnet 5, Opus 5, and Fable 5 already have model-level cost results in Labs. The reported 12×–40× Claude Code gap is a plan-packaging comparison, not a new model benchmark: Max allowances are usage-limited, Enterprise discounts are private, and the cited raw sessions are not a shared reproducible fixture. Labs will not publish a synthetic multiplier. A valid plan study needs consented per-call usage logs, identical Claude Code and model versions, cache-write/read splits, limit events, negotiated terms, and accepted-task outcomes across both routes. See the buyer-cost analysis and live calculator →

Rostered · Grok 4.6 result pending

The route is available; the funded refresh gate is not

Grok 4.6 and its matching x-ai/grok-4.6 OpenRouter route are in the maintained Labs roster, but no result is published yet. The benchmark account's funded refresh gate is blocked, so Labs preserves the last complete leaderboard rather than treating rejected calls, zero responses, or a price-only estimate as measured performance. The model will run through the same 49 tasks, deterministic graders, retry policy, and spend log after that gate clears. See the Grok 4.6 pricing and migration analysis →

Availability blocker · Grok Bot

A managed cloud teammate cannot be substituted with a Grok API call

Grok Bot is not ranked because xAI publishes no fixed model selector, API model ID, deterministic callable route, weekly allowance size, or per-Bot rate. Requests use a managed model set with automatic failover, while the product's value depends on browser and desktop actions, connectors, approvals, routines, and successful cross-app completion. Existing public Grok model results remain valid, but they do not measure Grok Bot. A reproducible study needs a stable test account, disclosed serving-model and usage logs, identical app fixtures, fixed approval rules, and cost per accepted end-to-end task. See the Grok Bot plan and billing analysis →

Task mix

classification: 6coding: 6extraction: 6instruction: 10pricing: 8reasoning: 8summarization: 5
View every benchmark prompt
  1. math-001 · reasoning

    Return only the integer answer. A batch job processes 18 files per minute for 7 minutes, then 12 files per minute for 5 minutes. How many files total?

  2. math-002 · reasoning

    Return only the integer answer. A model costs $0.30 per 1M input tokens and $2.50 per 1M output tokens. What is the cost in cents for 100,000 input tokens and 20,000 output tokens?

  3. math-003 · reasoning

    Return only the integer answer. If a 64,000-token context window is filled to 75%, how many tokens are used?

  4. math-004 · reasoning

    Return only the integer answer. A provider raises output price from $4/M to $6/M. What is the percent increase?

  5. math-005 · reasoning

    Return only the integer answer. A cache hit costs 10% of normal input price. Normal input is $2/M. What is the cache-hit price in cents per 1M tokens?

  6. math-006 · reasoning

    Return only the integer answer. A crawler finds 9 new models Monday, 14 Tuesday, and removes 5 duplicates. How many unique new models remain?

  7. math-007 · reasoning

    Return only the integer answer. A benchmark has 50 tasks. A model gets 37 correct. What is the accuracy percentage rounded to the nearest whole number?

  8. math-008 · reasoning

    Return only the decimal answer with no currency symbol. If a run spends $0.42 and gets 14 correct answers, what is dollars per correct answer?

  9. extract-001 · extraction

    Return only JSON with numeric values for prices and context, without currency symbols or units. Text: 'Gemini Flash: input $0.30/M, output $2.50/M, context 1,048,576 tokens.' Extract fields model, inputPerM, outputPerM, context.

  10. extract-002 · extraction

    Return only JSON. Text: 'Z.ai GLM-5.2 charges $1.40 per million input tokens and $4.40 per million output tokens.' Extract provider, model, inputPerM, outputPerM.

  11. extract-003 · extraction

    Return only the single word immediately before 'plan'. Sentence: 'For heavy coding, the Pro plan is cheaper than per-token API use after roughly 35 million output tokens.'

  12. extract-004 · extraction

    Return only JSON. Log line: '2026-07-20 model=deepseek-v4-flash prompt=120 completion=38 latency_ms=910 ok=true'. Extract model, prompt, completion, latency_ms, ok.

  13. extract-005 · extraction

    Return only a comma-separated list of model IDs in their original order. Text: 'Tested: gpt-5.6-luna; claude-sonnet-5; gemini-2.5-flash.'

  14. extract-006 · extraction

    Return only JSON. Text: 'Anthropic Claude Haiku 4.5: $1 input, $5 output, active.' Extract provider, model, status.

  15. classify-001 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'OpenAI cuts GPT-5.6 Luna output token price by 20%'.

  16. classify-002 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'DeepSeek introduces V4 Flash for low-latency chat apps'.

  17. classify-003 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'Claude API errors spike across us-east for 42 minutes'.

  18. classify-004 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'Step-by-step guide to add streaming responses to a chatbot'.

  19. classify-005 · classification

    Return only yes or no. Query: 'cheapest ai api per token'. Is this query primarily cost/comparison intent?

  20. classify-006 · classification

    Return only yes or no. Query: 'what is a transformer neural network'. Is this query primarily pricing intent?

  21. summ-001 · summarization

    Summarize in exactly 8 words and include these exact words: cheaper, extraction, failed, arithmetic. Text: 'The cheaper model answered all extraction tasks correctly but failed most multi-step arithmetic, making it a good fit for structured parsing rather than reasoning.'

  22. summ-002 · summarization

    Return these three search phrases exactly as written, lowercase and separated by commas with no extra text: ai api pricing, cheapest ai api, token calculator. Text: 'Rising searches include ai api pricing, cheapest ai api, and token calculator, all pointing toward cost-conscious developer intent.'

  23. summ-003 · summarization

    Return one sentence under 12 words and include the exact words verbosity and cost. Text: 'A model with high accuracy but very high output verbosity can lose on dollars per correct answer because every extra token compounds cost.'

  24. summ-004 · summarization

    Return only a title under 7 words that includes the exact words cost and correct. Text: 'We compare leading AI models by how much each correct benchmark answer costs, not just accuracy.'

  25. summ-005 · summarization

    Return exactly one sentence that includes OpenRouter and the exact phrase list prices. Text: 'OpenRouter is used for routing, but costs are computed from AI Pricing Guru's direct-provider list prices so the metric reflects what API buyers see on provider pricing pages.'

  26. code-001 · coding

    Return only the output. JavaScript: const xs=[3,1,4,1,5]; console.log(xs.filter(x=>x>2).reduce((a,b)=>a+b,0));

  27. code-002 · coding

    Return only the output. Python: prices={'in':0.3,'out':2.5}; print(round(prices['in']*2 + prices['out']*0.1, 2))

  28. code-003 · coding

    Return only the missing JavaScript expression. Complete: const cost = (inputTokens * ___ + outputTokens * outputPerM) / 1_000_000;

  29. code-004 · coding

    Return only JSON. For rows [{ok:true,cost:0.02},{ok:false,cost:0.03},{ok:true,cost:0.01}], compute correct, totalCost, costPerCorrect.

  30. code-005 · coding

    Return only the output. JavaScript: console.log(['gpt','claude','glm'].map(s=>s.length).join('-'));

  31. code-006 · coding

    Return only the called function name immediately after await. Snippet: async function call(){ const r = await fetch(url, opts); return r.json(); }

  32. logic-001 · instruction

    Return only the third item alphabetically: Luna, Flash, Sonnet, Grok.

  33. logic-002 · instruction

    Return only the word that appears twice: token, price, model, token, output.

  34. logic-003 · instruction

    Return only the reversed string: glm-5.2

  35. logic-004 · instruction

    Return only the uppercase acronym from: cost per correct answer.

  36. logic-005 · instruction

    Return only a comma-separated list of the two providers with names starting with G: OpenAI, Google, Anthropic, Groq, Mistral.

  37. logic-006 · instruction

    Return only valid or invalid. JSON string: {"model":"gpt","cost":0.01}

  38. logic-007 · instruction

    Return only valid or invalid. JSON string: {model:"gpt",cost:0.01}

  39. logic-008 · instruction

    Return only the smallest price: $1.25/M, $0.30/M, $2.00/M, $0.14/M.

  40. logic-009 · instruction

    Return only the final model ID. Chain: start=gpt-5; replace gpt with glm; append .2; prefix z-ai-.

  41. logic-010 · instruction

    Return only the count of unique providers: OpenAI, Anthropic, OpenAI, Google, Z.ai, Google.

  42. pricing-001 · pricing

    Return only cheaper or expensive. Model A costs $0.10/M input and $0.30/M output. Model B costs $0.30/M input and $2.50/M output. For any positive input/output workload, Model A is what relative to Model B?

  43. pricing-002 · pricing

    Return only the model name. A got 10 correct for $0.20. B got 8 correct for $0.08. Which has lower dollars per correct answer?

  44. pricing-003 · pricing

    Return only the integer answer. If output tokens are 4x more expensive than input tokens, how many input-token equivalents is 250 output tokens?

  45. pricing-004 · pricing

    Return only JSON. A run used 1000 prompt tokens and 500 completion tokens. Prices are $1/M input and $6/M output. Compute inputCost, outputCost, totalCost in dollars.

  46. pricing-005 · pricing

    Return only yes or no. If a model is free for input but charges for output, can a verbose wrong answer still cost money?

  47. pricing-006 · pricing

    Return only the ratio as N:1. Output costs $15/M and input costs $3/M.

  48. pricing-007 · pricing

    Return only the integer answer. A $5 pilot budget has already spent $1.75. How many cents remain?

  49. pricing-008 · pricing

    Return only the best metric name: cost per prompt, cost per correct answer, or total tokens. We need to compare accuracy and spend together.

AI Pricing Guru Labs is editorially independent. Affiliate relationships do not affect the models selected, prompts, grading, ranking, or reported results.

Coming next

More questions Labs will answer

Coding Plan Value

When does a fixed-price coding subscription beat paying per successful API task?

Long-context Cost

How does cost per correctly retrieved fact change as context grows?

Silent Nerf Watch

Does a model’s cost per correct answer worsen over time without a price change?