AI Earnings Forecast Benchmark: Pricing Impact
Samaya says frontier AI beat analyst consensus on earnings forecasts. Compare current API rates and the costs its benchmark leaves out.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- Samaya reports GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5 beat a bias-corrected analyst-consensus baseline on its earnings-forecast benchmark.
- This is a model-and-harness result, not a public API price change or proof that an off-the-shelf chatbot beats every analyst.
- Compare current direct API rates below; Samaya did not disclose cost per forecast or the price of its data and harness.
API cost of a comparable token workload
USD per 1M tokens. Input and output rates are charted separately.
Estimate your own forecast-agent token spend
Assumes 75% input tokens and 25% output tokens using current per-million rates.
Claude Sonnet 5
anthropic
$40.00
- Input share
- $15.00
- Output share
- $25.00
Claude Opus 5.5
anthropic
$80.00
- Input share
- $30.00
- Output share
- $50.00
GPT-6 Astra
openai
$200.00
- Input share
- $75.00
- Output share
- $125.00
Claude Fable 5.1
anthropic
$200.00
- Input share
- $75.00
- Output share
- $125.00
Current direct API rates for tested models—not Samaya forecast costs
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| GPT-6 Astra | openai | $10.00 | $1.00 | $50.00 |
| Claude Fable 5.1 | anthropic | $10.00 | $0.25 | $50.00 |
| Claude Opus 5.5 | anthropic | $4.00 | $0.2 | $20.00 |
| Claude Sonnet 5 | anthropic | $2.00 | $0.2 | $10.00 |
Built from pricing.json at publish time.
Samaya published an earnings-prediction benchmark on October 8, 2026. Its strongest finding is that GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5, inside Samaya’s finance-specific retrieval harness, outperformed a bias-corrected analyst-consensus baseline. That is a meaningful result for finance teams to test—not an announcement that any model’s API rate has changed.
What the benchmark actually measured
Samaya evaluated 456 companies reporting earnings from July 14 onward, each with a market capitalization above $5 billion and estimates from more than eight brokers. Five trading sessions before each release, its agents predicted revenue, gross margin, operating income, and adjusted earnings per share using information available at that point. The team scored forecast error, correlation with earnings surprises, and whether a forecast called a beat or miss relative to consensus.
The plain analyst-consensus baseline proved easier to beat: even Claude Sonnet 5 and Kimi K3 cleared it. Samaya therefore created a harder baseline by adjusting each company’s consensus for its historical median surprise. The three latest frontier models named above beat that stronger baseline in the company’s tests. Samaya says GPT-6 Astra led on overall error and hit rate, while Claude Fable 5.1 and Opus 5.5 performed best on surprise correlation. These are different measures; there is no single winner on every one.
| Test detail | Reported result | Buying implication |
|---|---|---|
| Prediction horizon | Five trading sessions before earnings | A same-day chatbot answer is not the same test |
| Coverage | 456 large, heavily covered companies | Do not extrapolate to thinly covered stocks |
| Stronger human baseline | Consensus adjusted for historical surprise | Compare against your best existing process |
| Best-performing routes | Astra, Fable 5.1, Opus 5.5 | Test accuracy and all-in cost per accepted forecast |
Pricing impact: public rates, not a benchmark bill
The live table and chart above compare direct model API list rates, not Samaya’s product price or the expense of reproducing its benchmark. No vendor announced a token-price cut in this research. The study does not give token counts, tool-call counts, data-licensing fees, retries, or billed cost per company. Its harness also supplies time-gated financial data and expert-guided instructions; a raw API prompt should not be expected to match the published result.
Samaya’s ablation found timely data especially important: moving from early-quarter data to information five sessions before earnings added a reported 12–16 percentage points of surprise-correlation performance. Expert guidance increased research time and final-call context use by roughly 1.6–2.7×. A more capable model can therefore raise the bill per forecast, even if it reduces forecast error.
For an internal pilot, use the token calculator with your measured input, cached-input, and output volumes. Then add market-data rights, retrieval infrastructure, analyst review, and failed or repeated runs. See the current OpenAI pricing and Anthropic pricing pages for model-specific billing rules; our OpenAI API pricing guide explains long-context and caching charges.
Samaya’s product-access page does not provide a public per-forecast rate. If you want a separately hosted Kimi K3 comparison, check Novita’s current Kimi K3 route and its exact terms; it is not Samaya’s licensed finance harness. Affiliate disclosure: we may earn a commission at no extra cost to you.
Our Labs coverage note records why we cannot import this result into our 49-task text leaderboard: the finance dataset, point-in-time gate, retrieval, prompts, raw traces, baseline and full bill are missing, and Astra is unavailable on our exact benchmark route.
What this means for finance teams
Potential beneficiaries: research desks with licensed point-in-time data and enough repeated forecasts to evaluate a controlled agent. A model that catches misses that analysts systematically overlook could justify a higher API rate—but only if those calls improve a decision or reduce review work.
Potential losers: teams that buy a frontier subscription or API access expecting Samaya’s result without the retrieval layer, historical consensus, and validation process. The study comes from the company selling the harness; it is not an independent head-to-head audit of standalone models or a demonstrated trading return.
Run a retrospective pilot that freezes every source at the prediction cutoff. Compare each model with raw consensus, a bias-corrected consensus, and your own analyst process on the same companies. Record forecast quality by metric, total tokens, cache hits, data cost, reviewer time, and cost per useful forecast. Keep a holdout set and check whether the advantage survives outside heavily covered large-cap names before expanding.
Bottom line: this research strengthens the case for testing frontier-model finance agents with current data. It does not establish a new API price, a guaranteed cheaper analyst workflow, or an investable edge on its own.
Source: Samaya’s October 8 research report. Direct model rates are supplied by our maintained canonical pricing feed, not this study. We checked Astra against OpenAI’s official API price card, and Fable 5.1 and Opus 5.5 against Anthropic’s Fable and Opus model pages on October 8, 2026.