AI Pricing Guru Labs · Experiment 01

Which AI model gives the most correct answers per dollar?

We ran 14 models through the same 49 short, machine-graded tasks. The primary score is cost per correct answer. A model must score at least 95% to qualify, so a low price cannot compensate for too many wrong answers.

See the verdict Compare all models Open prompts + raw results

Live experiment 01

One question, one primary metric

Question: for routine, short-form API work, which model produces correct answers at the lowest token cost? We divide each model’s total direct-price run cost by the number of correct answers. Lower is better—but only if the observed error rate is acceptable for your workflow.

Primary score: cost per correct answer. This combines the bill and the observed success rate into one number. Qualified models are sorted by this score, so there are no confusing accuracy ties.
Quality guardrail: at least 95% accuracy. A model needs 47/49 correct to receive a value rank. Lower-scoring models remain visible, but they are marked below threshold rather than recommended.
14 models in this run Task suite v2 Full benchmark 49 tasks per model 95% ranking threshold 686 answers graded Run 2026-07-23

The verdict

What works, what does not, and what it costs

Best value for this workload

DeepSeek V4 Pro

49/49 correct · 100% accuracy

$0.00658 for all 49 tasks · $0.000134 per correct answer. No misses across all 7 categories.

Cheapest zero-miss option

DeepSeek V4 Pro

49/49 correct · 100% accuracy

$0.00658 for the run · $0.000134 per correct answer. It bought 0 more correct answers than DeepSeek V4 Pro, but at 1× the cost per correct answer.

Poor value on these simple tasks

Gemini 3.6 Flash

49/49 correct · 100% accuracy

$0.094044 for the run · $0.001919 per correct answer, or 14.3× the best-value score. That does not make it a bad model; it means its premium was not justified by this workload.

Decision in plain English

Choose DeepSeek V4 Pro when occasional misses can be validated or retried. Choose DeepSeek V4 Pro when this exact task mix requires zero observed misses. At the same prompt lengths and mix, one million correct answers would cost roughly $134 versus $134. Do not extrapolate that estimate to longer or more complex workloads.

Cost-per-task leaderboard

Lowest cost per correct answer wins

Value rank #1

DeepSeek V4 Pro

$0.000134

per correct

Accuracy100%
Run cost$0.00658
Vs. bestBest

Perfect across all 7 categories. No misses in this run.

Value rank #2

GPT-5.6 Luna

$0.000262

per correct

Accuracy100%
Run cost$0.012822
Vs. best1.9×

Perfect across all 7 categories. No misses in this run.

Value rank #3

Claude Sonnet 5

$0.000447

per correct

Accuracy96%
Run cost$0.021004
Vs. best3.3×

Perfect: Reasoning, Extraction, Classification, Summarization, Coding
Missed: Instruction: 1 missed · Pricing: 1 missed

Value rank #4

GPT-5.6 Terra

$0.000534

per correct

Accuracy100%
Run cost$0.026175
Vs. best

Perfect across all 7 categories. No misses in this run.

Value rank #5

GLM-5.2

$0.000668

per correct

Accuracy100%
Run cost$0.032749
Vs. best

Perfect across all 7 categories. No misses in this run.

Value rank #6

GPT-5.6 Sol

$0.000822

per correct

Accuracy100%
Run cost$0.04026
Vs. best6.1×

Perfect across all 7 categories. No misses in this run.

Value rank #7

Claude Opus 4.8

$0.000905

per correct

Accuracy98%
Run cost$0.04346
Vs. best6.7×

Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Instruction
Missed: Pricing: 1 missed

Value rank #8

Grok 4.5

$0.0013

per correct

Accuracy100%
Run cost$0.063686
Vs. best9.7×

Perfect across all 7 categories. No misses in this run.

Value rank #9

Gemini 3.6 Flash

$0.001919

per correct

Accuracy100%
Run cost$0.094044
Vs. best14.3×

Perfect across all 7 categories. No misses in this run.

Below 95% threshold

DeepSeek V4 Flash

$0.000037

per correct

Accuracy90%
Run cost$0.001624
Vs. best0.3×

Perfect: Reasoning, Summarization, Pricing
Missed: Extraction: 1 missed · Classification: 1 missed · Coding: 2 missed · Instruction: 1 missed

Below 95% threshold

Mistral Large 3

$0.000054

per correct

Accuracy90%
Run cost$0.002383
Vs. best0.4×

Perfect: Extraction, Classification, Pricing
Missed: Reasoning: 2 missed · Summarization: 1 missed · Coding: 1 missed · Instruction: 1 missed

Below 95% threshold

Gemini 3.5 Flash-Lite

$0.00006

per correct

Accuracy94%
Run cost$0.002754
Vs. best0.4×

Perfect: Extraction, Classification, Coding, Pricing
Missed: Reasoning: 1 missed · Summarization: 1 missed · Instruction: 1 missed

Below 95% threshold

Claude Haiku 4.5

$0.000145

per correct

Accuracy94%
Run cost$0.006656
Vs. best1.1×

Perfect: Extraction, Classification, Summarization, Coding, Pricing
Missed: Reasoning: 2 missed · Instruction: 1 missed

Below 95% threshold

Claude Fable 5 Preview

$0.002877

per correct

Accuracy94%
Run cost$0.13232
Vs. best21.4×

Perfect: Extraction, Classification, Summarization, Instruction
Missed: Reasoning: 1 missed · Coding: 1 missed · Pricing: 1 missed

Models need at least 95% accuracy (47/49) to qualify. Qualified models are ranked by run cost ÷ correct answers. “Vs. best” shows how many times more each correct answer cost than the winner. Models below the quality threshold remain visible without receiving a value rank.

Methodology

Reproducible, deterministic grading

Same prompts and settings

Every model receives the same system prompt and user tasks through OpenRouter with temperature 0, top_p 1, and a 1024-token output ceiling so reasoning tokens do not truncate the final answer.

No LLM judge

Answers are checked with exact, numeric, regex, contains, and JSON-field graders. This avoids paying another model to make subjective scoring decisions.

Direct-provider price basis

OpenRouter is the inference route. Cost is calculated from the direct-provider list prices in AI Pricing Guru’s daily pricing dataset, using reported prompt and completion tokens.

What this does not prove

This compact suite does not measure long-context reliability, safety, creative writing, production latency, tool use, or every coding workload. Results are a workload signal, not a universal model ranking.

Task mix

classification: 6coding: 6extraction: 6instruction: 10pricing: 8reasoning: 8summarization: 5
View every benchmark prompt
  1. math-001 · reasoning

    Return only the integer answer. A batch job processes 18 files per minute for 7 minutes, then 12 files per minute for 5 minutes. How many files total?

  2. math-002 · reasoning

    Return only the integer answer. A model costs $0.30 per 1M input tokens and $2.50 per 1M output tokens. What is the cost in cents for 100,000 input tokens and 20,000 output tokens?

  3. math-003 · reasoning

    Return only the integer answer. If a 64,000-token context window is filled to 75%, how many tokens are used?

  4. math-004 · reasoning

    Return only the integer answer. A provider raises output price from $4/M to $6/M. What is the percent increase?

  5. math-005 · reasoning

    Return only the integer answer. A cache hit costs 10% of normal input price. Normal input is $2/M. What is the cache-hit price in cents per 1M tokens?

  6. math-006 · reasoning

    Return only the integer answer. A crawler finds 9 new models Monday, 14 Tuesday, and removes 5 duplicates. How many unique new models remain?

  7. math-007 · reasoning

    Return only the integer answer. A benchmark has 50 tasks. A model gets 37 correct. What is the accuracy percentage rounded to the nearest whole number?

  8. math-008 · reasoning

    Return only the decimal answer with no currency symbol. If a run spends $0.42 and gets 14 correct answers, what is dollars per correct answer?

  9. extract-001 · extraction

    Return only JSON with numeric values for prices and context, without currency symbols or units. Text: 'Gemini Flash: input $0.30/M, output $2.50/M, context 1,048,576 tokens.' Extract fields model, inputPerM, outputPerM, context.

  10. extract-002 · extraction

    Return only JSON. Text: 'Z.ai GLM-5.2 charges $1.40 per million input tokens and $4.40 per million output tokens.' Extract provider, model, inputPerM, outputPerM.

  11. extract-003 · extraction

    Return only the single word immediately before 'plan'. Sentence: 'For heavy coding, the Pro plan is cheaper than per-token API use after roughly 35 million output tokens.'

  12. extract-004 · extraction

    Return only JSON. Log line: '2026-07-20 model=deepseek-v4-flash prompt=120 completion=38 latency_ms=910 ok=true'. Extract model, prompt, completion, latency_ms, ok.

  13. extract-005 · extraction

    Return only a comma-separated list of model IDs in their original order. Text: 'Tested: gpt-5.6-luna; claude-sonnet-5; gemini-2.5-flash.'

  14. extract-006 · extraction

    Return only JSON. Text: 'Anthropic Claude Haiku 4.5: $1 input, $5 output, active.' Extract provider, model, status.

  15. classify-001 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'OpenAI cuts GPT-5.6 Luna output token price by 20%'.

  16. classify-002 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'DeepSeek introduces V4 Flash for low-latency chat apps'.

  17. classify-003 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'Claude API errors spike across us-east for 42 minutes'.

  18. classify-004 · classification

    Return only one label: pricing, launch, outage, or tutorial. Headline: 'Step-by-step guide to add streaming responses to a chatbot'.

  19. classify-005 · classification

    Return only yes or no. Query: 'cheapest ai api per token'. Is this query primarily cost/comparison intent?

  20. classify-006 · classification

    Return only yes or no. Query: 'what is a transformer neural network'. Is this query primarily pricing intent?

  21. summ-001 · summarization

    Summarize in exactly 8 words and include these exact words: cheaper, extraction, failed, arithmetic. Text: 'The cheaper model answered all extraction tasks correctly but failed most multi-step arithmetic, making it a good fit for structured parsing rather than reasoning.'

  22. summ-002 · summarization

    Return these three search phrases exactly as written, lowercase and separated by commas with no extra text: ai api pricing, cheapest ai api, token calculator. Text: 'Rising searches include ai api pricing, cheapest ai api, and token calculator, all pointing toward cost-conscious developer intent.'

  23. summ-003 · summarization

    Return one sentence under 12 words and include the exact words verbosity and cost. Text: 'A model with high accuracy but very high output verbosity can lose on dollars per correct answer because every extra token compounds cost.'

  24. summ-004 · summarization

    Return only a title under 7 words that includes the exact words cost and correct. Text: 'We compare leading AI models by how much each correct benchmark answer costs, not just accuracy.'

  25. summ-005 · summarization

    Return exactly one sentence that includes OpenRouter and the exact phrase list prices. Text: 'OpenRouter is used for routing, but costs are computed from AI Pricing Guru's direct-provider list prices so the metric reflects what API buyers see on provider pricing pages.'

  26. code-001 · coding

    Return only the output. JavaScript: const xs=[3,1,4,1,5]; console.log(xs.filter(x=>x>2).reduce((a,b)=>a+b,0));

  27. code-002 · coding

    Return only the output. Python: prices={'in':0.3,'out':2.5}; print(round(prices['in']*2 + prices['out']*0.1, 2))

  28. code-003 · coding

    Return only the missing JavaScript expression. Complete: const cost = (inputTokens * ___ + outputTokens * outputPerM) / 1_000_000;

  29. code-004 · coding

    Return only JSON. For rows [{ok:true,cost:0.02},{ok:false,cost:0.03},{ok:true,cost:0.01}], compute correct, totalCost, costPerCorrect.

  30. code-005 · coding

    Return only the output. JavaScript: console.log(['gpt','claude','glm'].map(s=>s.length).join('-'));

  31. code-006 · coding

    Return only the called function name immediately after await. Snippet: async function call(){ const r = await fetch(url, opts); return r.json(); }

  32. logic-001 · instruction

    Return only the third item alphabetically: Luna, Flash, Sonnet, Grok.

  33. logic-002 · instruction

    Return only the word that appears twice: token, price, model, token, output.

  34. logic-003 · instruction

    Return only the reversed string: glm-5.2

  35. logic-004 · instruction

    Return only the uppercase acronym from: cost per correct answer.

  36. logic-005 · instruction

    Return only a comma-separated list of the two providers with names starting with G: OpenAI, Google, Anthropic, Groq, Mistral.

  37. logic-006 · instruction

    Return only valid or invalid. JSON string: {"model":"gpt","cost":0.01}

  38. logic-007 · instruction

    Return only valid or invalid. JSON string: {model:"gpt",cost:0.01}

  39. logic-008 · instruction

    Return only the smallest price: $1.25/M, $0.30/M, $2.00/M, $0.14/M.

  40. logic-009 · instruction

    Return only the final model ID. Chain: start=gpt-5; replace gpt with glm; append .2; prefix z-ai-.

  41. logic-010 · instruction

    Return only the count of unique providers: OpenAI, Anthropic, OpenAI, Google, Z.ai, Google.

  42. pricing-001 · pricing

    Return only cheaper or expensive. Model A costs $0.10/M input and $0.30/M output. Model B costs $0.30/M input and $2.50/M output. For any positive input/output workload, Model A is what relative to Model B?

  43. pricing-002 · pricing

    Return only the model name. A got 10 correct for $0.20. B got 8 correct for $0.08. Which has lower dollars per correct answer?

  44. pricing-003 · pricing

    Return only the integer answer. If output tokens are 4x more expensive than input tokens, how many input-token equivalents is 250 output tokens?

  45. pricing-004 · pricing

    Return only JSON. A run used 1000 prompt tokens and 500 completion tokens. Prices are $1/M input and $6/M output. Compute inputCost, outputCost, totalCost in dollars.

  46. pricing-005 · pricing

    Return only yes or no. If a model is free for input but charges for output, can a verbose wrong answer still cost money?

  47. pricing-006 · pricing

    Return only the ratio as N:1. Output costs $15/M and input costs $3/M.

  48. pricing-007 · pricing

    Return only the integer answer. A $5 pilot budget has already spent $1.75. How many cents remain?

  49. pricing-008 · pricing

    Return only the best metric name: cost per prompt, cost per correct answer, or total tokens. We need to compare accuracy and spend together.

AI Pricing Guru Labs is editorially independent. Affiliate relationships do not affect the models selected, prompts, grading, ranking, or reported results.

Coming next

More questions Labs will answer

Coding Plan Value

When does a fixed-price coding subscription beat paying per successful API task?

Long-context Cost

How does cost per correctly retrieved fact change as context grows?

Silent Nerf Watch

Does a model’s cost per correct answer worsen over time without a price change?