Spotify Portal Cuts Claude Code Context 90%: Cost Impact
Spotify reports 90% less Claude context with Portal shunt. The benchmark estimates tokens and shifts work to Gemini—here’s the real cost impact.
By AI Pricing Guru Editorial Team
AI Pricing Guru articles are maintained by the editorial workflow behind the site: daily pricing snapshots, provider source checks, and review passes for model launches, subscription limits, and billing changes.
TL;DR
- Spotify reports 82–94% less Claude context across three large-file reads, with a 90% mean, after its shunt plugin routed the files through a Gemini worker.
- The benchmark estimates tokens as characters divided by four; it does not use provider token ledgers, measure the worker's full token usage, or prove a 90% lower total bill.
- The plugin is Apache-2.0, but it requires a Portal instance with AiKA and a configured worker model; Spotify does not publish an all-in Portal price on the cited pages.
- Best use: large-file summaries and boilerplate. Keep debugging, architecture, safety-critical code, and final acceptance with the stronger model and a human reviewer.
Current API cost for the same fixed token workload
USD per 1M tokens. Input and output rates are charted separately.
Model your Claude-to-Gemini routing mix
Assumes 75% input tokens and 25% output tokens using current per-million rates.
Gemini 2.5 Flash
$8.50
- Input share
- $2.25
- Output share
- $6.25
Gemini 3.8 Flash
$15.00
- Input share
- $5.63
- Output share
- $9.38
Claude Sonnet 5
anthropic
$40.00
- Input share
- $15.00
- Output share
- $25.00
Live model rates—the worker still consumes tokens
| Model | Provider | Input / 1M | Cached / 1M | Output / 1M |
|---|---|---|---|---|
| Claude Sonnet 5 | anthropic | $2.00 | $0.2 | $10.00 |
| Gemini 2.5 Flash | $0.3 | $0.03 | $2.50 | |
| Gemini 3.8 Flash | $0.75 | $0.075 | $3.75 |
Built from pricing.json at publish time.
Spotify says its shunt plugin reduced the context sent back to Claude Code by an average of 90% across three large-file reading scenarios. The mechanism is model routing: hooks stop Claude from opening large files in full, send those files to a Portal AiKA mode using Gemini 2.5 Flash, and return a shorter summary to Claude.
That is a potentially useful reduction in frontier-model context, not a measured 90% reduction in total tokens or total spend. The public benchmark estimates tokens from character counts, and the files still pass through a worker model. No Anthropic or Google rate changed with the September 3 post.
What Spotify released
shunt is a Claude Code plugin in Spotify’s public portal-ai-plugins repository. A PreToolUse hook blocks full reads above 350 lines by default and points Claude to a bulk-reader skill. A second hook catches large reads through commands such as cat, head, and tail; targeted reads and piped searches can pass.
The bulk-reader calls an ephemeral AiKA mode through the Portal CLI. A separate code-writer mode can generate predictable files directly on disk so Claude does not receive the generated output in its context. Spotify uses Gemini 2.5 Flash in its example, but Portal administrators can configure a different worker.
The repository is Apache-2.0 licensed. Operation is not necessarily free: users need a Portal instance with AiKA, authentication, and a configured model route. Spotify’s public “Try Portal” page offers a walkthrough but does not show an all-in list price, so procurement cost must be added to model usage and setup time.
What the 90% benchmark measured
Spotify’s README publishes results from a 162,000-line Java monorepo. The three bulk-read rows average to 90%; the separate code-writing row is not included in that mean.
| Scenario | Lines | Estimated Claude context without shunt | Estimated Claude context with shunt | Reported saving |
|---|---|---|---|---|
| Single large file | 4,014 | 33,684 tokens | 5,737 tokens | 82% |
| Source and test pair | 7,408 | 75,990 tokens | 4,148 tokens | 94% |
| Cross-service files | 1,281 | 16,221 tokens | 821 tokens | 94% |
The benchmark script defines a token estimate as characters divided by four. For bulk reads, “without” is the estimated file content and “with” is the estimated AiKA summary returned to Claude. It does not report Anthropic or Google billing records, the worker’s input and output totals, repeated trials, correctness scores, or accepted-change rates.
Spotify also reports that its worker missed a subtle thread-safety bug that Claude found once given the relevant context. That limitation matters more than the headline: a short summary can save context while dropping the evidence needed for a correct decision.
Pricing impact: shifted tokens, not vanished tokens
The live table and chart above pull rates from our maintained dataset. Gemini 2.5 Flash is shown because it powered Spotify’s example, but it is now a legacy model in our registry; Gemini 3.8 Flash is included as the current Flash reference. A team reproducing the setup should verify that its Portal route still supports the chosen worker.
There are at least four budget lines:
| Cost layer | Effect of shunt |
|---|---|
| Claude context | Falls when a summary replaces full files |
| Worker-model input and output | Added for every delegated call |
| Portal/AiKA | Commercial access and operating terms are not priced in the benchmark |
| Engineering quality | Latency, review, retries, and missed context can erase token savings |
API users can compare the two model bills directly. Claude subscription users may instead experience fewer usage-limit events without seeing a proportional cash saving. In both cases, optimize cost per accepted task, not Claude-visible tokens alone.
Who should test it—and who should not
The strongest fit is repetitive, high-volume I/O: summarizing several large files, extracting symbols, matching an established test pattern, or generating configuration boilerplate. Spotify says each delegation adds roughly 10–30 seconds, so small files are a poor target.
Do not route debugging, architecture, safety-critical work, or edits that require exact line-level context through a lossy summary by default. The plugin itself allows targeted reads for editing, and its code-writer path relies on Claude choosing the skill rather than a blocking hook.
For a production test, pin the repository commit, Claude Code version, worker model, threshold, and task set. Record both providers’ input, cache, and output tokens; latency; retries; tests passed; human review time; and whether the change was accepted. Compare at least 20 paired tasks before changing team-wide defaults.
For another context-reduction approach, see our Graft Claude Code token analysis. Compare current rates on the Anthropic pricing page and Google AI pricing page, then test your actual mix with the token calculator.
Teams evaluating a separately metered coding route can also compare the Z.ai coding plan on the same accepted-task benchmark.
Affiliate disclosure: AI Pricing Guru may earn a commission from the sponsored Z.ai link at no extra cost to you. It does not affect this analysis.
Bottom line
Spotify’s shunt plugin is a concrete implementation of a sensible idea: reserve the stronger model for judgment and send bulk I/O to a cheaper worker. The published 82–94% reduction in Claude-visible context is large enough to justify a test.
It is not yet evidence of 90% lower total token consumption or spend. The benchmark uses a character-based estimate, omits worker usage and quality scoring, and sits inside a Portal deployment with undisclosed all-in pricing. Reproduce the result on accepted work before putting 90% into a budget forecast.
Sources: Spotify Engineering’s Portal and Claude Code report, Spotify’s shunt plugin and benchmark, the Portal trial page, and the Hacker News discussion. Sources and prices checked September 5, 2026.