Affiliate disclosure: we may earn commissions when you sign up through some links below, at no extra cost to you. This never affects our pricing data, comparisons, or recommendations. Learn more.

Alibaba Qwen price tracker

Qwen API Pricing 2026

Compare first-party Alibaba Model Studio aliases with hosted Qwen routes from Cerebras, Novita, Together, Castform, and Groq. The tables use the same daily-refreshed structured data as our main API comparison, including cache and long-context fields where the provider publishes them.

33 current routes across 5 providers · Rates refreshed · Qwen Studio plan status checked

First-party Alibaba aliases

4

Managed Qwen routes

29

Lowest tracked input rate

$0.03

Qwen3.5 4B

Alibaba Model Studio Qwen prices

These are direct first-party international Model Studio aliases in USD per 1 million tokens. Current promotions are temporary and have no published end date, so the acquisition gate checks the alias, list rate, promotional label, tier boundary, and rate every day.

Relative priceLowestLowerHigherHighestAutomatic log scale across current offers; input, cache and output are graded separately. Color shows cost, not quality.
Showing 4 grouped models from 4 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

New hosted route · Cerebras

Qwen3.8-27B runs at about 1,500 tokens/s for $0.99/$1.49

Cerebras now exposes qwen-3.8-27b on its public endpoints. Developer pricing is $0.99 per 1M input tokens and $1.49 per 1M output tokens. The paid route has 128K context, a 40K maximum output, and published speed of about 1,500 output tokens per second.

The hosted route supports text and image input, reasoning, streaming, tool calling, structured outputs, and prompt caching. Throughput is a provider measurement rather than an end-to-end latency guarantee, so benchmark time to first token and accepted-task cost on the exact route.

Read the Cerebras pricing, speed, and local-cost analysis →

Announced price · API coming soon

Qwen3.8-Flash targets $0.16 input and $0.47 output

Qwen opened the 125B-parameter Qwen3.8-Flash-Next weights and announced a production qwen3.8-flash QwenCloud SKU with 1M context. The official launch page publishes $0.16 per 1M input tokens and $0.47 per 1M output tokens, but it also says the API is not live yet. The structured row therefore stays marked preview / coming soon until the callable route appears in an official catalog.

The open weights use the Qwen Community License 1.0, not Apache 2.0. Its commercial-hosting and AI work-assistant conditions deserve legal review before self-hosted redistribution or a managed service launch.

Read the pricing, architecture, and availability analysis →

New local route · Junie Local

JetBrains packages Qwen3.6-27B for free on M5 Macs

Junie Local installs a JetBrains-tuned 4-bit Qwen3.6-27B agent with one command and meters no tokens, credits, or subscription. The initial release requires an Apple Silicon M5 Mac, 64 GB RAM, macOS later than version 26, and about 40 GB of free disk; the model-weights download is roughly 20 GB. It is a local product route, not a new Alibaba or Novita API SKU, so the canonical token-pricing rows stay unchanged.

JetBrains reports about 40% higher prefill throughput from its M5-specific patch and roughly 2× faster generation from speculative decoding, but those are vendor measurements. Compare accepted repository changes, wall time, power, retries, and hardware depreciation against the managed Qwen3.6-27B row below.

Read the Junie Local pricing and Mac requirements analysis →

New self-hosting benchmark · Qwen3-TTS

Nari reports 34 ms p95 first audio at 10 RPS on one H100

Nari Labs open-sourced an Apache-2.0 serving engine for Qwen3-TTS 1.7B CustomVoice. This is not a hosted Nari API or new Alibaba SKU, so no row is added to the token-pricing tables. At Nari's reported 630 characters per second, Lambda's current $4.29 one-GPU H100 SXM hourly price implies a GPU-only floor of about $1.89 per 1M characters at full utilization; idle capacity, networking, operations, redundancy, and quality testing are extra.

Read the benchmark audit and verified cost math →

Request-size pricing matters

Plus and Flash become more expensive above 256K input tokens

The headline table shows the standard tier. The structured rows for Qwen3.7-Plus and Qwen3.6-Flash also include the provider's higher tier for requests above 256K input tokens. The Qwen3.7-Max alias currently uses one promotional tier through its 1M context window.

Model your token mix in the calculator →

Managed Qwen API prices

A managed route is a separate product and bill. It may expose open Qwen weights, an Alibaba closed-model route, or a provider-specific snapshot with different caching, context, throughput, and support. Compare exact IDs rather than assuming two rows with similar names are interchangeable.

Relative priceLowestLowerHigherHighestAutomatic log scale across current offers; input, cache and output are graded separately. Color shows cost, not quality.
Showing 29 grouped models from 29 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

Is Qwen Studio free?

Qwen's public Studio experience does not currently expose a stable monthly subscription rate card or fixed public quota that we can maintain as a plan. That is why Qwen is not inserted into the subscription matrix as an invented zero-dollar plan. The separate API has published token billing and selected activation quotas.

See Qwen Studio plan and limit status →

Want one endpoint for Qwen and DeepSeek?

Novita is an active AI Pricing Guru partner and carries multiple Qwen routes alongside DeepSeek and other open-model families. Verify the exact model ID and run the same acceptance test before choosing a host.

Check Novita's current catalog →

Affiliate disclosure: we may earn a commission at no extra cost to you. Partner status does not change the table or ordering.

Which Qwen model should you choose?

Start with Flash for inexpensive classification, extraction, retrieval, and short structured responses. Move to Plus when the task needs stronger reasoning, coding, or long-context work. Use Max as an escalation route only when it produces enough fewer failures or revisions to justify the higher accepted-task cost.

For coding-specific open models, compare the hosted Qwen Coder rows in the managed table and read the Qwen Code pricing guide. For buyer-oriented price and capability tradeoffs, see Qwen vs DeepSeek API pricing.

Open weights are not the same as a free API

Many Qwen models can be downloaded under permissive model licenses, but production self-hosting still pays for accelerators, memory, electricity or rental, redundancy, monitoring, and engineering time. Compare those costs in the local AI versus API calculator before treating a downloadable checkpoint as free inference.

Frequently asked questions

How much does the Qwen API cost?

Qwen pricing depends on the exact model, request size, cache use, region, and host. This page tracks 33 current or announced Qwen routes across 5 providers. Cerebras now serves Qwen3.8-27B at $0.99 input and $1.49 output per 1 million tokens.

Is Qwen free?

Qwen Studio provides a consumer chat experience, but its public site does not publish a stable paid-plan rate card or fixed public usage allowance. Alibaba Model Studio separately offers activation quotas for selected international API models; after the applicable quota, API usage is billed per token. Do not treat free chat access as a free production API.

What is the cheapest hosted Qwen model?

Qwen3.5 4B is the lowest current tracked Qwen route by standard input rate at $0.03 per 1 million input tokens. Output price, retries, context tier, and host reliability can change the cheapest accepted-task result.

Should I use Alibaba Model Studio or a managed Qwen host?

Use Alibaba Model Studio when you want first-party aliases, regional controls, and Qwen-specific features. Use a managed host when one compatible endpoint, cross-family routing, or an existing cloud relationship matters more. Benchmark the exact snapshot because model names, cache rules, and prices can differ by host.

Does Qwen long-context pricing cost more?

Some first-party Qwen aliases change rates when the request crosses a documented input-token threshold. The Qwen3.7-Plus and Qwen3.6-Flash rows on this page carry their higher long-context tiers in the live structured dataset rather than hiding the higher rate in prose.

Sources and maintenance

First-party models are checked against Alibaba Cloud's model catalog, international pricing page, and cache rules. The Cerebras route is checked against its model page and pricing page. Qwen Studio status is checked against Qwen's official product. All are supervised on a daily cadence.