Alibaba Qwen price tracker
Qwen API Pricing 2026
Compare first-party Alibaba Model Studio aliases with hosted Qwen routes from Cerebras, Novita, Together, Castform, and Groq. The tables use the same daily-refreshed structured data as our main API comparison, including cache and long-context fields where the provider publishes them.
33 current routes across 5 providers · Rates refreshed · Qwen Studio plan status checked
First-party Alibaba aliases
4
Managed Qwen routes
29
Lowest tracked input rate
$0.03
Qwen3.5 4B
Alibaba Model Studio Qwen prices
These are direct first-party international Model Studio aliases in USD per 1 million tokens. Current promotions are temporary and have no published end date, so the acquisition gate checks the alias, list rate, promotional label, tier boundary, and rate every day.
| Try it | |||||||
|---|---|---|---|---|---|---|---|
Qwen3.7-Max | Alibaba | Mid | 1M | $1.25 | $0.125 | $3.75 | Try API → |
Qwen3.8-FlashPreview | Alibaba | Low | 1M | $0.16 | - | $0.47 | Try API → |
Qwen3.6-Flash | Alibaba | Low | 1M | $0.25 | $0.025 | $1.50 | Try API → |
Qwen3.7-Plus | Alibaba | Low | 1M | $0.32 | $0.032 | $1.28 | Try API → |
- Qwen3.7-MaxAlibabaMid
- Input
- $1.25
- Cached
- $0.125
- Output
- $3.75
- Qwen3.8-FlashPreviewAlibabaLow
- Input
- $0.16
- Cached
- -
- Output
- $0.47
- Qwen3.6-FlashAlibabaLow
- Input
- $0.25
- Cached
- $0.025
- Output
- $1.50
- Qwen3.7-PlusAlibabaLow
- Input
- $0.32
- Cached
- $0.032
- Output
- $1.28
Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.
All product names, logos, and brands are property of their respective owners and are used for identification purposes only.
New hosted route · Cerebras
Qwen3.8-27B runs at about 1,500 tokens/s for $0.99/$1.49
Cerebras now exposes qwen-3.8-27b on its public endpoints. Developer pricing is $0.99 per 1M input tokens and $1.49 per 1M output tokens. The paid route has 128K context, a 40K maximum output, and published speed of about 1,500 output tokens per second.
The hosted route supports text and image input, reasoning, streaming, tool calling, structured outputs, and prompt caching. Throughput is a provider measurement rather than an end-to-end latency guarantee, so benchmark time to first token and accepted-task cost on the exact route.
Read the Cerebras pricing, speed, and local-cost analysis →Announced price · API coming soon
Qwen3.8-Flash targets $0.16 input and $0.47 output
Qwen opened the 125B-parameter Qwen3.8-Flash-Next weights and announced a production qwen3.8-flash QwenCloud SKU with 1M context. The official launch page publishes $0.16 per 1M input tokens and $0.47 per 1M output tokens, but it also says the API is not live yet. The structured row therefore stays marked preview / coming soon until the callable route appears in an official catalog.
The open weights use the Qwen Community License 1.0, not Apache 2.0. Its commercial-hosting and AI work-assistant conditions deserve legal review before self-hosted redistribution or a managed service launch.
Read the pricing, architecture, and availability analysis →New local route · Junie Local
JetBrains packages Qwen3.6-27B for free on M5 Macs
Junie Local installs a JetBrains-tuned 4-bit Qwen3.6-27B agent with one command and meters no tokens, credits, or subscription. The initial release requires an Apple Silicon M5 Mac, 64 GB RAM, macOS later than version 26, and about 40 GB of free disk; the model-weights download is roughly 20 GB. It is a local product route, not a new Alibaba or Novita API SKU, so the canonical token-pricing rows stay unchanged.
JetBrains reports about 40% higher prefill throughput from its M5-specific patch and roughly 2× faster generation from speculative decoding, but those are vendor measurements. Compare accepted repository changes, wall time, power, retries, and hardware depreciation against the managed Qwen3.6-27B row below.
Read the Junie Local pricing and Mac requirements analysis →New self-hosting benchmark · Qwen3-TTS
Nari reports 34 ms p95 first audio at 10 RPS on one H100
Nari Labs open-sourced an Apache-2.0 serving engine for Qwen3-TTS 1.7B CustomVoice. This is not a hosted Nari API or new Alibaba SKU, so no row is added to the token-pricing tables. At Nari's reported 630 characters per second, Lambda's current $4.29 one-GPU H100 SXM hourly price implies a GPU-only floor of about $1.89 per 1M characters at full utilization; idle capacity, networking, operations, redundancy, and quality testing are extra.
Read the benchmark audit and verified cost math →Request-size pricing matters
Plus and Flash become more expensive above 256K input tokens
The headline table shows the standard tier. The structured rows for Qwen3.7-Plus and Qwen3.6-Flash also include the provider's higher tier for requests above 256K input tokens. The Qwen3.7-Max alias currently uses one promotional tier through its 1M context window.
Model your token mix in the calculator →Managed Qwen API prices
A managed route is a separate product and bill. It may expose open Qwen weights, an Alibaba closed-model route, or a provider-specific snapshot with different caching, context, throughput, and support. Compare exact IDs rather than assuming two rows with similar names are interchangeable.
| Try it | |||||||
|---|---|---|---|---|---|---|---|
Qwen3.5 9B | Together | Low | - | $0.17 | - | $0.25 | Try API → |
Qwen3 235B A22B Instruct 2507 | Together | Low | - | $0.20 | - | $0.60 | Try API → |
Qwen2.5 7B Instruct Turbo | Together | Low | - | $0.30 | - | $0.30 | Try API → |
Qwen3.5 397B A17B | Together | Low | - | $0.60 | $0.35 | $3.60 | Try API → |
Qwen3.7-Max | Novita | Mid | - | $1.25 | $0.25 | $3.75 | Try on Novita → |
Qwen3.8 2.4T A95B | Novita | Mid | - | $2.00 | $0.25 | $6.00 | Try on Novita → |
Qwen3.8 Max | Novita | Mid | - | $2.00 | $0.25 | $6.00 | Try on Novita → |
CQwen3.5 4B | Ccastform | Low | - | $0.03 | - | $0.15 | Open provider → |
Qwen3 Coder 30B A3B Instruct | Novita | Low | - | $0.07 | - | $0.27 | Try on Novita → |
Qwen3 235B A22B Instruct 2507 | Novita | Low | - | $0.09 | - | $0.58 | Try on Novita → |
Qwen3 Next 80B A3B Instruct | Novita | Low | - | $0.15 | - | $1.50 | Try on Novita → |
Qwen3.8 Flash | Novita | Low | - | $0.15 | $0.016 | $0.47 | Try on Novita → |
Qwen3 235B A22B | Novita | Low | - | $0.20 | - | $0.80 | Try on Novita → |
Qwen3 Coder Next | Novita | Low | - | $0.20 | - | $1.50 | Try on Novita → |
Qwen3 VL 30B A3B Instruct | Novita | Low | - | $0.20 | - | $0.70 | Try on Novita → |
Qwen3.6-35B-A3B | Novita | Low | - | $0.248 | - | $1.49 | Try on Novita → |
Qwen MT Plus | Novita | Low | - | $0.25 | - | $0.75 | Try on Novita → |
Qwen3.5-35B-A3B | Novita | Low | - | $0.25 | - | $2.00 | Try on Novita → |
Qwen3 235B A22B Thinking 2507 | Novita | Low | - | $0.30 | - | $3.00 | Try on Novita → |
Qwen3 VL 235B A22B Instruct | Novita | Low | - | $0.30 | - | $1.50 | Try on Novita → |
Qwen3.5-27B | Novita | Low | - | $0.30 | - | $2.40 | Try on Novita → |
Qwen 2.5 72B Instruct | Novita | Low | - | $0.38 | - | $0.40 | Try on Novita → |
Qwen3 Coder 480B A35B Instruct | Novita | Low | - | $0.38 | - | $1.55 | Try on Novita → |
Qwen3.5-122B-A10B | Novita | Low | - | $0.40 | - | $3.20 | Try on Novita → |
Qwen3.8 27B | Novita | Low | - | $0.42 | $0.085 | $3.00 | Try on Novita → |
Qwen3.5-397B-A17B | Novita | Low | - | $0.60 | - | $3.60 | Try on Novita → |
Qwen3.6-27B | Novita | Low | - | $0.60 | - | $3.60 | Try on Novita → |
Qwen3 VL 235B A22B Thinking | Novita | Low | - | $0.98 | - | $3.95 | Try on Novita → |
CQwen3.8 27B | Ccerebras | Low | 131,072 | $0.99 | - | $1.49 | — |
- Qwen3.5 9BTogetherLow
- Input
- $0.17
- Cached
- -
- Output
- $0.25
- Qwen3 235B A22B Instruct 2507TogetherLow
- Input
- $0.20
- Cached
- -
- Output
- $0.60
- Qwen2.5 7B Instruct TurboTogetherLow
- Input
- $0.30
- Cached
- -
- Output
- $0.30
- Qwen3.5 397B A17BTogetherLow
- Input
- $0.60
- Cached
- $0.35
- Output
- $3.60
- Qwen3.7-MaxNovitaMid
- Input
- $1.25
- Cached
- $0.25
- Output
- $3.75
- Qwen3.8 2.4T A95BNovitaMid
- Input
- $2.00
- Cached
- $0.25
- Output
- $6.00
- Qwen3.8 MaxNovitaMid
- Input
- $2.00
- Cached
- $0.25
- Output
- $6.00
- CQwen3.5 4BCcastformLow
- Input
- $0.03
- Cached
- -
- Output
- $0.15
- Qwen3 Coder 30B A3B InstructNovitaLow
- Input
- $0.07
- Cached
- -
- Output
- $0.27
- Qwen3 235B A22B Instruct 2507NovitaLow
- Input
- $0.09
- Cached
- -
- Output
- $0.58
- Qwen3 Next 80B A3B InstructNovitaLow
- Input
- $0.15
- Cached
- -
- Output
- $1.50
- Qwen3.8 FlashNovitaLow
- Input
- $0.15
- Cached
- $0.016
- Output
- $0.47
- Qwen3 235B A22BNovitaLow
- Input
- $0.20
- Cached
- -
- Output
- $0.80
- Qwen3 Coder NextNovitaLow
- Input
- $0.20
- Cached
- -
- Output
- $1.50
- Qwen3 VL 30B A3B InstructNovitaLow
- Input
- $0.20
- Cached
- -
- Output
- $0.70
- Qwen3.6-35B-A3BNovitaLow
- Input
- $0.248
- Cached
- -
- Output
- $1.49
- Qwen MT PlusNovitaLow
- Input
- $0.25
- Cached
- -
- Output
- $0.75
- Qwen3.5-35B-A3BNovitaLow
- Input
- $0.25
- Cached
- -
- Output
- $2.00
- Qwen3 235B A22B Thinking 2507NovitaLow
- Input
- $0.30
- Cached
- -
- Output
- $3.00
- Qwen3 VL 235B A22B InstructNovitaLow
- Input
- $0.30
- Cached
- -
- Output
- $1.50
- Qwen3.5-27BNovitaLow
- Input
- $0.30
- Cached
- -
- Output
- $2.40
- Qwen 2.5 72B InstructNovitaLow
- Input
- $0.38
- Cached
- -
- Output
- $0.40
- Qwen3 Coder 480B A35B InstructNovitaLow
- Input
- $0.38
- Cached
- -
- Output
- $1.55
- Qwen3.5-122B-A10BNovitaLow
- Input
- $0.40
- Cached
- -
- Output
- $3.20
- Qwen3.8 27BNovitaLow
- Input
- $0.42
- Cached
- $0.085
- Output
- $3.00
- Qwen3.5-397B-A17BNovitaLow
- Input
- $0.60
- Cached
- -
- Output
- $3.60
- Qwen3.6-27BNovitaLow
- Input
- $0.60
- Cached
- -
- Output
- $3.60
- Qwen3 VL 235B A22B ThinkingNovitaLow
- Input
- $0.98
- Cached
- -
- Output
- $3.95
- CQwen3.8 27BCcerebrasLow
- Input
- $0.99
- Cached
- -
- Output
- $1.49
—
Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.
All product names, logos, and brands are property of their respective owners and are used for identification purposes only.
Is Qwen Studio free?
Qwen's public Studio experience does not currently expose a stable monthly subscription rate card or fixed public quota that we can maintain as a plan. That is why Qwen is not inserted into the subscription matrix as an invented zero-dollar plan. The separate API has published token billing and selected activation quotas.
See Qwen Studio plan and limit status →Want one endpoint for Qwen and DeepSeek?
Novita is an active AI Pricing Guru partner and carries multiple Qwen routes alongside DeepSeek and other open-model families. Verify the exact model ID and run the same acceptance test before choosing a host.
Check Novita's current catalog →Affiliate disclosure: we may earn a commission at no extra cost to you. Partner status does not change the table or ordering.
Which Qwen model should you choose?
Start with Flash for inexpensive classification, extraction, retrieval, and short structured responses. Move to Plus when the task needs stronger reasoning, coding, or long-context work. Use Max as an escalation route only when it produces enough fewer failures or revisions to justify the higher accepted-task cost.
For coding-specific open models, compare the hosted Qwen Coder rows in the managed table and read the Qwen Code pricing guide. For buyer-oriented price and capability tradeoffs, see Qwen vs DeepSeek API pricing.
Open weights are not the same as a free API
Many Qwen models can be downloaded under permissive model licenses, but production self-hosting still pays for accelerators, memory, electricity or rental, redundancy, monitoring, and engineering time. Compare those costs in the local AI versus API calculator before treating a downloadable checkpoint as free inference.
Frequently asked questions
How much does the Qwen API cost?
Qwen pricing depends on the exact model, request size, cache use, region, and host. This page tracks 33 current or announced Qwen routes across 5 providers. Cerebras now serves Qwen3.8-27B at $0.99 input and $1.49 output per 1 million tokens.
Is Qwen free?
Qwen Studio provides a consumer chat experience, but its public site does not publish a stable paid-plan rate card or fixed public usage allowance. Alibaba Model Studio separately offers activation quotas for selected international API models; after the applicable quota, API usage is billed per token. Do not treat free chat access as a free production API.
What is the cheapest hosted Qwen model?
Qwen3.5 4B is the lowest current tracked Qwen route by standard input rate at $0.03 per 1 million input tokens. Output price, retries, context tier, and host reliability can change the cheapest accepted-task result.
Should I use Alibaba Model Studio or a managed Qwen host?
Use Alibaba Model Studio when you want first-party aliases, regional controls, and Qwen-specific features. Use a managed host when one compatible endpoint, cross-family routing, or an existing cloud relationship matters more. Benchmark the exact snapshot because model names, cache rules, and prices can differ by host.
Does Qwen long-context pricing cost more?
Some first-party Qwen aliases change rates when the request crosses a documented input-token threshold. The Qwen3.7-Plus and Qwen3.6-Flash rows on this page carry their higher long-context tiers in the live structured dataset rather than hiding the higher rate in prose.
Sources and maintenance
First-party models are checked against Alibaba Cloud's model catalog, international pricing page, and cache rules. The Cerebras route is checked against its model page and pricing page. Qwen Studio status is checked against Qwen's official product. All are supervised on a daily cadence.