AI Pricing Guru Labs

Benchmark coverage notes

Not every model announcement, vendor benchmark, or product claim belongs in one deterministic text leaderboard. These decisions document what is measured, what stays separate, and what evidence would make a future comparison valid.

30 documented decisions Reviewed 2026-09-03 Return to the leaderboard

The inclusion rule

A result enters the main leaderboard only when the exact model or route can run the same published tasks with fixed settings, machine-graded outputs, complete token records, and a defensible current price basis. The linked analyses contain the evidence and current pricing context; the benchmark itself remains vendor-neutral.

Coverage notes are not benchmark scores. When a model becomes callable or a reproducible artifact appears, we rerun or reprice the relevant experiment and update the main Labs leaderboard.