Best value for this workload
DeepSeek V4 Pro
49/49 correct · 100% accuracy
$0.00658 for all 49 tasks · $0.000134 per correct answer. No misses across all 7 categories.
AI Pricing Guru Labs · Experiment 01
We ran 14 models through the same 49 short, machine-graded tasks. The primary score is cost per correct answer. A model must score at least 95% to qualify, so a low price cannot compensate for too many wrong answers.
Live experiment 01
Question: for routine, short-form API work, which model produces correct answers at the lowest token cost? We divide each model’s total direct-price run cost by the number of correct answers. Lower is better—but only if the observed error rate is acceptable for your workflow.
The verdict
Best value for this workload
49/49 correct · 100% accuracy
$0.00658 for all 49 tasks · $0.000134 per correct answer. No misses across all 7 categories.
Cheapest zero-miss option
49/49 correct · 100% accuracy
$0.00658 for the run · $0.000134 per correct answer. It bought 0 more correct answers than DeepSeek V4 Pro, but at 1× the cost per correct answer.
Poor value on these simple tasks
49/49 correct · 100% accuracy
$0.094044 for the run · $0.001919 per correct answer, or 14.3× the best-value score. That does not make it a bad model; it means its premium was not justified by this workload.
Choose DeepSeek V4 Pro when occasional misses can be validated or retried. Choose DeepSeek V4 Pro when this exact task mix requires zero observed misses. At the same prompt lengths and mix, one million correct answers would cost roughly $134 versus $134. Do not extrapolate that estimate to longer or more complex workloads.
Cost-per-task leaderboard
Value rank #1
$0.000134
per correct
Perfect across all 7 categories. No misses in this run.
Value rank #2
$0.000262
per correct
Perfect across all 7 categories. No misses in this run.
Value rank #3
$0.000447
per correct
Perfect: Reasoning, Extraction, Classification, Summarization, Coding
Missed: Instruction: 1 missed · Pricing: 1 missed
Value rank #4
$0.000534
per correct
Perfect across all 7 categories. No misses in this run.
Value rank #5
$0.000668
per correct
Perfect across all 7 categories. No misses in this run.
Value rank #6
$0.000822
per correct
Perfect across all 7 categories. No misses in this run.
Value rank #7
$0.000905
per correct
Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Instruction
Missed: Pricing: 1 missed
Value rank #8
$0.0013
per correct
Perfect across all 7 categories. No misses in this run.
Value rank #9
$0.001919
per correct
Perfect across all 7 categories. No misses in this run.
Below 95% threshold
$0.000037
per correct
Perfect: Reasoning, Summarization, Pricing
Missed: Extraction: 1 missed · Classification: 1 missed · Coding: 2 missed · Instruction: 1 missed
Below 95% threshold
$0.000054
per correct
Perfect: Extraction, Classification, Pricing
Missed: Reasoning: 2 missed · Summarization: 1 missed · Coding: 1 missed · Instruction: 1 missed
Below 95% threshold
$0.00006
per correct
Perfect: Extraction, Classification, Coding, Pricing
Missed: Reasoning: 1 missed · Summarization: 1 missed · Instruction: 1 missed
Below 95% threshold
$0.000145
per correct
Perfect: Extraction, Classification, Summarization, Coding, Pricing
Missed: Reasoning: 2 missed · Instruction: 1 missed
Below 95% threshold
$0.002877
per correct
Perfect: Extraction, Classification, Summarization, Instruction
Missed: Reasoning: 1 missed · Coding: 1 missed · Pricing: 1 missed
| Value rank | Model | Accuracy | What worked / missed | Run cost | Cost / correct | Vs. best |
|---|---|---|---|---|---|---|
| #1 | DeepSeek V4 Pro | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.00658 | $0.000134 | Best |
| #2 | GPT-5.6 Luna | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.012822 | $0.000262 | 1.9× |
| #3 | Claude Sonnet 5 | 96% 47/49 | Perfect: Reasoning, Extraction, Classification, Summarization, Coding Missed: Instruction: 1 missed · Pricing: 1 missed | $0.021004 | $0.000447 | 3.3× |
| #4 | GPT-5.6 Terra | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.026175 | $0.000534 | 4× |
| #5 | GLM-5.2 | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.032749 | $0.000668 | 5× |
| #6 | GPT-5.6 Sol | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.04026 | $0.000822 | 6.1× |
| #7 | Claude Opus 4.8 | 98% 48/49 | Perfect: Reasoning, Extraction, Classification, Summarization, Coding, Instruction Missed: Pricing: 1 missed | $0.04346 | $0.000905 | 6.7× |
| #8 | Grok 4.5 | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.063686 | $0.0013 | 9.7× |
| #9 | Gemini 3.6 Flash | 100% 49/49 | Perfect across all 7 categories. No misses in this run. | $0.094044 | $0.001919 | 14.3× |
| Below 95% | DeepSeek V4 Flash | 90% 44/49 | Perfect: Reasoning, Summarization, Pricing Missed: Extraction: 1 missed · Classification: 1 missed · Coding: 2 missed · Instruction: 1 missed | $0.001624 | $0.000037 | 0.3× |
| Below 95% | Mistral Large 3 | 90% 44/49 | Perfect: Extraction, Classification, Pricing Missed: Reasoning: 2 missed · Summarization: 1 missed · Coding: 1 missed · Instruction: 1 missed | $0.002383 | $0.000054 | 0.4× |
| Below 95% | Gemini 3.5 Flash-Lite | 94% 46/49 | Perfect: Extraction, Classification, Coding, Pricing Missed: Reasoning: 1 missed · Summarization: 1 missed · Instruction: 1 missed | $0.002754 | $0.00006 | 0.4× |
| Below 95% | Claude Haiku 4.5 | 94% 46/49 | Perfect: Extraction, Classification, Summarization, Coding, Pricing Missed: Reasoning: 2 missed · Instruction: 1 missed | $0.006656 | $0.000145 | 1.1× |
| Below 95% | Claude Fable 5 Preview | 94% 46/49 | Perfect: Extraction, Classification, Summarization, Instruction Missed: Reasoning: 1 missed · Coding: 1 missed · Pricing: 1 missed | $0.13232 | $0.002877 | 21.4× |
Models need at least 95% accuracy (47/49) to qualify. Qualified models are ranked by run cost ÷ correct answers. “Vs. best” shows how many times more each correct answer cost than the winner. Models below the quality threshold remain visible without receiving a value rank.
Methodology
Every model receives the same system prompt and user tasks through OpenRouter with temperature 0, top_p 1, and a 1024-token output ceiling so reasoning tokens do not truncate the final answer.
Answers are checked with exact, numeric, regex, contains, and JSON-field graders. This avoids paying another model to make subjective scoring decisions.
OpenRouter is the inference route. Cost is calculated from the direct-provider list prices in AI Pricing Guru’s daily pricing dataset, using reported prompt and completion tokens.
This compact suite does not measure long-context reliability, safety, creative writing, production latency, tool use, or every coding workload. Results are a workload signal, not a universal model ranking.
Return only the integer answer. A batch job processes 18 files per minute for 7 minutes, then 12 files per minute for 5 minutes. How many files total?
Return only the integer answer. A model costs $0.30 per 1M input tokens and $2.50 per 1M output tokens. What is the cost in cents for 100,000 input tokens and 20,000 output tokens?
Return only the integer answer. If a 64,000-token context window is filled to 75%, how many tokens are used?
Return only the integer answer. A provider raises output price from $4/M to $6/M. What is the percent increase?
Return only the integer answer. A cache hit costs 10% of normal input price. Normal input is $2/M. What is the cache-hit price in cents per 1M tokens?
Return only the integer answer. A crawler finds 9 new models Monday, 14 Tuesday, and removes 5 duplicates. How many unique new models remain?
Return only the integer answer. A benchmark has 50 tasks. A model gets 37 correct. What is the accuracy percentage rounded to the nearest whole number?
Return only the decimal answer with no currency symbol. If a run spends $0.42 and gets 14 correct answers, what is dollars per correct answer?
Return only JSON with numeric values for prices and context, without currency symbols or units. Text: 'Gemini Flash: input $0.30/M, output $2.50/M, context 1,048,576 tokens.' Extract fields model, inputPerM, outputPerM, context.
Return only JSON. Text: 'Z.ai GLM-5.2 charges $1.40 per million input tokens and $4.40 per million output tokens.' Extract provider, model, inputPerM, outputPerM.
Return only the single word immediately before 'plan'. Sentence: 'For heavy coding, the Pro plan is cheaper than per-token API use after roughly 35 million output tokens.'
Return only JSON. Log line: '2026-07-20 model=deepseek-v4-flash prompt=120 completion=38 latency_ms=910 ok=true'. Extract model, prompt, completion, latency_ms, ok.
Return only a comma-separated list of model IDs in their original order. Text: 'Tested: gpt-5.6-luna; claude-sonnet-5; gemini-2.5-flash.'
Return only JSON. Text: 'Anthropic Claude Haiku 4.5: $1 input, $5 output, active.' Extract provider, model, status.
Return only one label: pricing, launch, outage, or tutorial. Headline: 'OpenAI cuts GPT-5.6 Luna output token price by 20%'.
Return only one label: pricing, launch, outage, or tutorial. Headline: 'DeepSeek introduces V4 Flash for low-latency chat apps'.
Return only one label: pricing, launch, outage, or tutorial. Headline: 'Claude API errors spike across us-east for 42 minutes'.
Return only one label: pricing, launch, outage, or tutorial. Headline: 'Step-by-step guide to add streaming responses to a chatbot'.
Return only yes or no. Query: 'cheapest ai api per token'. Is this query primarily cost/comparison intent?
Return only yes or no. Query: 'what is a transformer neural network'. Is this query primarily pricing intent?
Summarize in exactly 8 words and include these exact words: cheaper, extraction, failed, arithmetic. Text: 'The cheaper model answered all extraction tasks correctly but failed most multi-step arithmetic, making it a good fit for structured parsing rather than reasoning.'
Return these three search phrases exactly as written, lowercase and separated by commas with no extra text: ai api pricing, cheapest ai api, token calculator. Text: 'Rising searches include ai api pricing, cheapest ai api, and token calculator, all pointing toward cost-conscious developer intent.'
Return one sentence under 12 words and include the exact words verbosity and cost. Text: 'A model with high accuracy but very high output verbosity can lose on dollars per correct answer because every extra token compounds cost.'
Return only a title under 7 words that includes the exact words cost and correct. Text: 'We compare leading AI models by how much each correct benchmark answer costs, not just accuracy.'
Return exactly one sentence that includes OpenRouter and the exact phrase list prices. Text: 'OpenRouter is used for routing, but costs are computed from AI Pricing Guru's direct-provider list prices so the metric reflects what API buyers see on provider pricing pages.'
Return only the output. JavaScript: const xs=[3,1,4,1,5]; console.log(xs.filter(x=>x>2).reduce((a,b)=>a+b,0));
Return only the output. Python: prices={'in':0.3,'out':2.5}; print(round(prices['in']*2 + prices['out']*0.1, 2))
Return only the missing JavaScript expression. Complete: const cost = (inputTokens * ___ + outputTokens * outputPerM) / 1_000_000;
Return only JSON. For rows [{ok:true,cost:0.02},{ok:false,cost:0.03},{ok:true,cost:0.01}], compute correct, totalCost, costPerCorrect.
Return only the output. JavaScript: console.log(['gpt','claude','glm'].map(s=>s.length).join('-'));
Return only the called function name immediately after await. Snippet: async function call(){ const r = await fetch(url, opts); return r.json(); }
Return only the third item alphabetically: Luna, Flash, Sonnet, Grok.
Return only the word that appears twice: token, price, model, token, output.
Return only the reversed string: glm-5.2
Return only the uppercase acronym from: cost per correct answer.
Return only a comma-separated list of the two providers with names starting with G: OpenAI, Google, Anthropic, Groq, Mistral.
Return only valid or invalid. JSON string: {"model":"gpt","cost":0.01}
Return only valid or invalid. JSON string: {model:"gpt",cost:0.01}
Return only the smallest price: $1.25/M, $0.30/M, $2.00/M, $0.14/M.
Return only the final model ID. Chain: start=gpt-5; replace gpt with glm; append .2; prefix z-ai-.
Return only the count of unique providers: OpenAI, Anthropic, OpenAI, Google, Z.ai, Google.
Return only cheaper or expensive. Model A costs $0.10/M input and $0.30/M output. Model B costs $0.30/M input and $2.50/M output. For any positive input/output workload, Model A is what relative to Model B?
Return only the model name. A got 10 correct for $0.20. B got 8 correct for $0.08. Which has lower dollars per correct answer?
Return only the integer answer. If output tokens are 4x more expensive than input tokens, how many input-token equivalents is 250 output tokens?
Return only JSON. A run used 1000 prompt tokens and 500 completion tokens. Prices are $1/M input and $6/M output. Compute inputCost, outputCost, totalCost in dollars.
Return only yes or no. If a model is free for input but charges for output, can a verbose wrong answer still cost money?
Return only the ratio as N:1. Output costs $15/M and input costs $3/M.
Return only the integer answer. A $5 pilot budget has already spent $1.75. How many cents remain?
Return only the best metric name: cost per prompt, cost per correct answer, or total tokens. We need to compare accuracy and spend together.
Coming next
When does a fixed-price coding subscription beat paying per successful API task?
How does cost per correctly retrieved fact change as context grows?
Does a model’s cost per correct answer worsen over time without a price change?