Meta Llama API Pricing 2026 — Live Hosted Rates

Last checked:

Meta does not publish one first-party Llama API rate. The table below renders 10 active hosted Llama routes from 3 providers in our daily canonical dataset. The lowest tracked input rate is $0.020 per 1M tokens on Novita, while the lowest output rate is $0.050 on Novita.

Current hosted Llama token prices

Relative priceLowestLowerHigherHighestAutomatic log scale across current offers; input, cache and output are graded separately. Color shows cost, not quality.
Showing 10 grouped models from 10 offers · USD per 1M tokens
Last synced:

Sponsored links may earn us a commission at no extra cost to you. Affiliate status never changes model ordering.

All product names, logos, and brands are property of their respective owners and are used for identification purposes only.

Rates are USD per 1 million tokens. Provider rows are verified by the recurring pricing pipeline; legacy routes are excluded from this current comparison.

August 10 release watch

Meta promised more open-source models—but did not release one today

Meta says it will “resume releasing some open source models soon.” The official announcement names no model, checkpoint, license, hardware target, public API, or token price. We therefore added no fabricated model or zero-dollar API row; the hosted Llama routes above remain the current cost baseline. Read the pricing-impact analysis →

How to compare hosted Llama with self-hosting

Match the exact checkpoint first. Then compare token rates, context limits, quantization, throughput, first-token latency, rate limits, support, and regional availability. A cheaper row can cost more per accepted task if it produces more retries or misses your latency target.

Self-hosting changes the unit of account from tokens to infrastructure. Include accelerator rental or depreciation, idle capacity, autoscaling, storage, networking, observability, upgrades, and operator time. Use our self-hosted AI cost calculator and API versus GPU break-even guide before assuming downloaded weights mean zero cost.

Labs availability blocker

Meta's announced future model cannot enter the Cost-per-Task Labs leaderboard yet. There is no named checkpoint, versioned artifact, callable route, measured usage, or official price. Substituting an older Llama model would test a different product. Labs will evaluate the release only after an exact artifact and reproducible serving route exist.

Price history

Only hosted routes with a recorded price change are charted here.

Llama 3.3 70B

Frequently asked questions

Did Meta release a new open-weight model on August 10, 2026?

No. Meta said it will resume releasing some open-source models soon, but it did not publish a model name, checkpoint, license, hardware target, callable endpoint, or price. Existing hosted Llama models remain the only priceable Meta-model baseline.

Does Meta sell a first-party Llama API?

Meta publishes Llama weights and documentation, but the active token prices tracked here belong to third-party inference hosts. A downloadable checkpoint does not establish a first-party Meta token price.

Why can the same Llama model have different prices?

The host controls serving hardware, batching, latency, support, rate limits, and margin. Compare the exact checkpoint and quantization as well as input and output rates before routing production traffic.

Is an open-weight Llama model free to run?

Open weights can remove a first-party per-token license toll, but inference still consumes GPUs, storage, power, observability, networking, and engineering time. Managed hosts package those costs into their rates.

Is Meta's next open-source model included in Labs?

Not yet. There is no exact artifact or supported inference route to test. Labs will require a named checkpoint, reproducible serving configuration, measurable usage, and a compatible evaluation route.

Methodology

Hosted prices render from the same daily-supervised canonical dataset published through our pricing API. The August 10 status is checked against Meta's official announcement. Open weights never create a zero-dollar hosted API row unless an official provider publishes that service and rate.