Included model · migration replay boundary
AI Pricing Guru Labs
Benchmark coverage notes
Not every model announcement, vendor benchmark, or product claim belongs in one deterministic text leaderboard. These decisions document what is measured, what stays separate, and what evidence would make a future comparison valid.
The inclusion rule
A result enters the main leaderboard only when the exact model or route can run the same published tasks with fixed settings, machine-graded outputs, complete token records, and a defensible current price basis. The linked analyses contain the evidence and current pricing context; the benchmark itself remains vendor-neutral.
Policy scope blocker · Mistral data training
A model benchmark cannot verify provider-side training use
Read the evidence and buyer impact →Included · GPT-5.6 Sol direct price cut
Fresh 49-task run uses OpenAI's new promotional rate
Read the evidence and buyer impact →Image-harness blocker · Grok Imagine retirement
The November 2 redirect cannot be scored in a text suite
Read the evidence and buyer impact →Router and training blocker · Experiential
A traffic-fitted gateway is not one model route
Read the evidence and buyer impact →Availability and physical-harness blocker · Model Hardware Standard
MHS is a hardware specification, not a model route
Read the evidence and buyer impact →Included models with creative-writing replay blocker · Gemini, Claude, OpenAI
Student preference is outside the deterministic 49-task score
Read the evidence and buyer impact →Included with Kiro-harness blocker · GPT-5.6 family
The models are ranked; Kiro's 82% claim is not reproduced
Read the evidence and buyer impact →Local-harness blocker · Junie Local Qwen3.6-27B
The managed base model is available; JetBrains' local package is not comparable here
Read the evidence and buyer impact →Included with route-price boundary · GPT-5.6 Sol
Sol stays ranked at the canonical direct rate
Read the evidence and buyer impact →Included with visual-harness blocker · GPT-5.6 Sol
Sol is ranked for text; Roboflow's image scores stay separate
Read the evidence and buyer impact →Included · DeepSeek V4 peak/off-peak pricing
Existing results stay valid; cost recalculation waits for the effective rate
Read the evidence and buyer impact →Included with evidence boundary · GPT-5.6 builder guidance
The models stay ranked; provider-run harness claims stay separate
Read the evidence and buyer impact →Security-harness blocker · Claude Code Auto mode
Opus 5 stays ranked; a targeted bypass cannot become a text-benchmark score
Read the evidence and buyer impact →Speech-harness blocker · Nari Qwen3-TTS
The artifacts are open; the current leaderboard measures the wrong task
Read the evidence and buyer impact →Harness blocker · Vomit for Claude Code
Claude is ranked; the local rewrite layer needs paired fidelity testing
Read the evidence and buyer impact →Harness blocker · Graft for Claude Code
The model is ranked; the context layer needs a paired repository study
Read the evidence and buyer impact →Availability blocker · Z.ai Ox Alpha
The preview route is live; its final identity and durable price are not
Read the evidence and buyer impact →Availability blocker · Qwen3.8-Flash
The price is official; the endpoint is not live yet
Read the evidence and buyer impact →Route blocker · Qwen3.8-27B on Cerebras
The API price is live; the 1,500 tokens/s claim awaits a comparable Labs run
Read the evidence and buyer impact →Service-tier blocker · GPT-5.6 Sol Ultrafast
The model is already ranked; the preview tier is not reproducible yet
Read the evidence and buyer impact →Controlled-route blocker · GPT-6 Astra
The price is live; initial access and ARC-AGI-3 harnesses are not yet comparable
Read the evidence and buyer impact →Availability blocker · GPT-5.6 Cyber
Daybreak Red is controlled access, not a public benchmark route
Read the evidence and buyer impact →Availability blocker · GPT-Realtime-2.1
Realtime voice needs a different benchmark harness
Read the evidence and buyer impact →Availability blocker · Meta's next open-source model
There is no model artifact to benchmark yet
Read the evidence and buyer impact →Evaluation blocker · Claude content marking
The announced detection API is not available for a reproducible test
Read the evidence and buyer impact →Scope blocker · Claude output-training policy
A terms boundary is not a model-performance result
Read the evidence and buyer impact →Scope blocker · Claude Code plan economics
Labs can price the models, but not reproduce a private seat allowance
Read the evidence and buyer impact →Rostered · Grok 4.6 result pending
The route is available; the funded refresh gate is not
Read the evidence and buyer impact →Availability blocker · Grok Bot