How Many AI Tokens Can I Get for My Budget?
Enter a monthly dollar amount and instantly see how many tokens each AI model gives you. Last updated:
How far does my AI budget go? Different models have wildly different per-token rates. A $50 monthly budget buys you about 2,500 million input tokens on Llama 3.1 8B Instruct but only about 2 million on GPT-5.4 Pro. This calculator divides your budget by each model's input and output rates and ranks every model by total token allowance, so you can find the best bang for your buck.
Your Budget
Token Allowance per Model — $50.00/mo Budget
Sorted by total tokens (most to least). Split: 50% input / 50% output.
| # | Model | Provider | Input Tokens | Output Tokens | Total Tokens |
|---|---|---|---|---|---|
| 1 | Embed v3 English | cohere | 250.0M | InfinityB | InfinityB |
| 2 | Embed v3 Multilingual | cohere | 250.0M | InfinityB | InfinityB |
| 3 | Rerank v3 | cohere | 12.5M | InfinityB | InfinityB |
| 4 | Llama 3.1 8B Instruct | novita | 1.3B | 500.0M | 1.8B |
| 5 | DeepSeek-OCR 2 | novita | 833.3M | 833.3M | 1.7B |
| 6 | LFM2 24B A2B | together | 833.3M | 208.3M | 1.0B |
| 7 | Qwen3.5 4B | castform | 833.3M | 166.7M | 1.0B |
| 8 | AutoGLM-Phone-9B-Multilingual | novita | 714.3M | 181.2M | 895.4M |
| 9 | Command R7B | cohere | 666.7M | 166.7M | 833.3M |
| 10 | Llama 3.1 8B Instant | groq | 500.0M | 312.5M | 812.5M |
| 11 | Mistral NeMo | novita | 625.0M | 147.1M | 772.1M |
| 12 | Gemma 3n E4B Instruct | together | 416.7M | 208.3M | 625.0M |
| 13 | GPT-OSS 20B | together | 500.0M | 125.0M | 625.0M |
| 14 | GPT-5 nano | openai | 500.0M | 62.5M | 562.5M |
| 15 | Ministral 3B | mistral | 250.0M | 250.0M | 500.0M |
| 16 | Qwen3 Coder 30B A3B Instruct | novita | 357.1M | 92.6M | 449.7M |
| 17 | GLM 5.3 Flash | novita | 333.3M | 100.0M | 433.3M |
| 18 | GLM-4.7-Flash | novita | 357.1M | 62.5M | 419.6M |
| 19 | Gemini 2.0 Flash-Lite | 333.3M | 83.3M | 416.7M | |
| 20 | GPT-OSS 20B | groq | 333.3M | 83.3M | 416.7M |
| 21 | GPT-OSS Safeguard 20B | groq | 333.3M | 83.3M | 416.7M |
| 22 | Llama 3 8B Instruct Lite | together | 178.6M | 178.6M | 357.1M |
| 23 | Devstral Small 2 | mistral | 250.0M | 83.3M | 333.3M |
| 24 | Llama 4 Scout | together | 250.0M | 83.3M | 333.3M |
| 25 | Ministral 8B | mistral | 166.7M | 166.7M | 333.3M |
| 26 | Mistral NeMo | mistral | 166.7M | 166.7M | 333.3M |
| 27 | Pixtral 12B | mistral | 166.7M | 166.7M | 333.3M |
| 28 | RNJ-1 Instruct | together | 166.7M | 166.7M | 333.3M |
| 29 | Qwen3 235B A22B Instruct 2507 | novita | 277.8M | 43.1M | 320.9M |
| 30 | Gemini 2.0 Flash | 250.0M | 62.5M | 312.5M | |
| 31 | Gemini 2.5 Flash-Lite | 250.0M | 62.5M | 312.5M | |
| 32 | GPT-4.1 nano | openai | 250.0M | 62.5M | 312.5M |
| 33 | Llama 4 Scout 17B 16E Instruct | groq | 227.3M | 73.5M | 300.8M |
| 34 | DeepSeek V4 Flash | novita | 178.6M | 89.3M | 267.9M |
| 35 | Ministral 14B | mistral | 125.0M | 125.0M | 250.0M |
| 36 | Llama 3.3 70B Instruct | novita | 185.2M | 62.5M | 247.7M |
| 37 | Qwen3.5 9B | together | 147.1M | 100.0M | 247.1M |
| 38 | GLM-4.5 Air | novita | 192.3M | 29.4M | 221.7M |
| 39 | Qwen3.8 Flash | novita | 166.7M | 53.2M | 219.9M |
| 40 | Qwen3.8-Flash | alibaba | 156.3M | 53.2M | 209.4M |
| 41 | Command R 08-2024 | cohere | 166.7M | 41.7M | 208.3M |
| 42 | GPT-4o mini | openai | 166.7M | 41.7M | 208.3M |
| 43 | GPT-OSS 120B | groq | 166.7M | 41.7M | 208.3M |
| 44 | GPT-OSS 120B | together | 166.7M | 41.7M | 208.3M |
| 45 | Llama 4 Maverick | together | 166.7M | 41.7M | 208.3M |
| 46 | Mistral Small 4 | mistral | 166.7M | 41.7M | 208.3M |
| 47 | Mistral 7B | mistral | 100.0M | 100.0M | 200.0M |
| 48 | Qwen3 Next 80B A3B Instruct | novita | 166.7M | 16.7M | 183.3M |
| 49 | Llama 4 Scout Instruct | novita | 138.9M | 42.4M | 181.3M |
| 50 | Grok 4.1 Fast Non-Reasoning | xai | 125.0M | 50.0M | 175.0M |
| 51 | Grok 4.1 Fast Reasoning | xai | 125.0M | 50.0M | 175.0M |
| 52 | Qwen2.5 7B Instruct Turbo | together | 83.3M | 83.3M | 166.7M |
| 53 | Qwen3 235B A22B Instruct 2507 | together | 125.0M | 41.7M | 166.7M |
| 54 | Qwen3 VL 30B A3B Instruct | novita | 125.0M | 35.7M | 160.7M |
| 55 | Qwen3 235B A22B | novita | 125.0M | 31.3M | 156.3M |
| 56 | DeepSeek V3.2 | novita | 92.9M | 62.5M | 155.4M |
| 57 | DeepSeek V3.2 Exp | novita | 92.6M | 61.0M | 153.6M |
| 58 | DeepSeek V4 Flash 0731 | deepseek | 113.6M | 37.9M | 151.5M |
| 59 | GPT-5.6 Luna | openai | 125.0M | 20.8M | 145.8M |
| 60 | GPT-5.4 nano | openai | 125.0M | 20.0M | 145.0M |
| 61 | Qwen3 Coder Next | novita | 125.0M | 16.7M | 141.7M |
| 62 | Qwen MT Plus | novita | 100.0M | 33.3M | 133.3M |
| 63 | Qwen3 32B | groq | 86.2M | 42.4M | 128.6M |
| 64 | Qwen 2.5 72B Instruct | novita | 65.8M | 62.5M | 128.3M |
| 65 | Llama 4 Maverick Instruct | novita | 92.6M | 29.4M | 122.0M |
| 66 | Gemma 4 31B IT Pearl | together | 89.3M | 29.1M | 118.4M |
| 67 | Qwen3.6-35B-A3B | novita | 100.8M | 16.8M | 117.6M |
| 68 | DeepSeek V3.1 | novita | 92.6M | 25.0M | 117.6M |
| 69 | DeepSeek V3.1 Terminus | novita | 92.6M | 25.0M | 117.6M |
| 70 | Gemini 3.1 Flash-Lite | 100.0M | 16.7M | 116.7M | |
| 71 | Qwen3.6-Flash | alibaba | 100.0M | 16.7M | 116.7M |
| 72 | DeepSeek V3 0324 | novita | 92.6M | 22.3M | 114.9M |
| 73 | GPT-5 mini | openai | 100.0M | 12.5M | 112.5M |
| 74 | Qwen3.5-35B-A3B | novita | 100.0M | 12.5M | 112.5M |
| 75 | Codestral | mistral | 83.3M | 27.8M | 111.1M |
| 76 | GLM-4.6V | novita | 83.3M | 27.8M | 111.1M |
| 77 | MiniMax M2 | novita | 83.3M | 20.8M | 104.2M |
| 78 | MiniMax M2.1 | novita | 83.3M | 20.8M | 104.2M |
| 79 | MiniMax M2.5 | together | 83.3M | 20.8M | 104.2M |
| 80 | MiniMax M2.5 | novita | 83.3M | 20.8M | 104.2M |
| 81 | MiniMax M2.7 | together | 83.3M | 20.8M | 104.2M |
| 82 | MiniMax M2.7 | novita | 83.3M | 20.8M | 104.2M |
| 83 | MiniMax M3 | minimax | 83.3M | 20.8M | 104.2M |
| 84 | MiniMax M3 | together | 83.3M | 20.8M | 104.2M |
| 85 | Qwen3 VL 235B A22B Instruct | novita | 83.3M | 16.7M | 100.0M |
| 86 | Qwen3.7-Plus | alibaba | 78.1M | 19.5M | 97.7M |
| 87 | Qwen3.5-27B | novita | 83.3M | 10.4M | 93.8M |
| 88 | Gemini 2.5 Flash | 83.3M | 10.0M | 93.3M | |
| 89 | Gemini 3.5 Flash-Lite | 83.3M | 10.0M | 93.3M | |
| 90 | Qwen3 235B A22B Thinking 2507 | novita | 83.3M | 8.3M | 91.7M |
| 91 | Gemma 4 31B IT | together | 64.1M | 25.8M | 89.9M |
| 92 | Qwen3 Coder 480B A35B Instruct | novita | 65.8M | 16.1M | 81.9M |
| 93 | DeepSeek V3 (Turbo) | novita | 62.5M | 19.2M | 81.7M |
| 94 | GPT-4.1 mini | openai | 62.5M | 15.6M | 78.1M |
| 95 | DeepSeek V4 Flash 0731 | novita | 56.8M | 18.9M | 75.8M |
| 96 | Devstral Medium 2 | mistral | 62.5M | 12.5M | 75.0M |
| 97 | Mistral Medium 3 | mistral | 62.5M | 12.5M | 75.0M |
| 98 | Llama 3.3 70B Versatile | groq | 42.4M | 31.6M | 74.0M |
| 99 | Mixtral 8x7B | mistral | 35.7M | 35.7M | 71.4M |
| 100 | Qwen3.5-122B-A10B | novita | 62.5M | 7.8M | 70.3M |
| 101 | Qwen3.8 27B | novita | 59.5M | 8.3M | 67.9M |
| 102 | gpt-3.5-turbo | openai | 50.0M | 16.7M | 66.7M |
| 103 | gpt-3.5-turbo-0125 | openai | 50.0M | 16.7M | 66.7M |
| 104 | Magistral Small | mistral | 50.0M | 16.7M | 66.7M |
| 105 | Mistral Large 3 | mistral | 50.0M | 16.7M | 66.7M |
| 106 | DeepSeek R1 Distill Llama 70B | novita | 31.3M | 31.3M | 62.5M |
| 107 | Gemini 3 Flash | 50.0M | 8.3M | 58.3M | |
| 108 | GLM-4.6 | novita | 45.5M | 11.4M | 56.8M |
| 109 | MiniMax M1 | novita | 45.5M | 11.4M | 56.8M |
| 110 | DeepSeek V3.1 | together | 41.7M | 14.7M | 56.4M |
| 111 | GLM-4.5V | novita | 41.7M | 13.9M | 55.6M |
| 112 | Kimi K2 Instruct | novita | 43.9M | 10.9M | 54.7M |
| 113 | GLM-4.7 | novita | 41.7M | 11.4M | 53.0M |
| 114 | MiniMax M2.5 Highspeed | novita | 41.7M | 10.4M | 52.1M |
| 115 | Kimi K2 0905 | novita | 41.7M | 10.0M | 51.7M |
| 116 | Kimi K2 Thinking | novita | 41.7M | 10.0M | 51.7M |
| 117 | DeepSeek V4 Pro 0813 | deepseek | 37.9M | 12.6M | 50.5M |
| 118 | Kimi K2.5 | novita | 41.7M | 8.3M | 50.0M |
| 119 | Sonar | perplexity | 25.0M | 25.0M | 50.0M |
| 120 | Nemotron 3 Ultra 550B A55B | together | 41.7M | 6.9M | 48.6M |
| 121 | Qwen3.5 397B A17B | together | 41.7M | 6.9M | 48.6M |
| 122 | Qwen3.5-397B-A17B | novita | 41.7M | 6.9M | 48.6M |
| 123 | Qwen3.6-27B | novita | 41.7M | 6.9M | 48.6M |
| 124 | Llama 3.3 70B | together | 24.0M | 24.0M | 48.1M |
| 125 | DeepSeek R1 (Turbo) | novita | 35.7M | 10.0M | 45.7M |
| 126 | DeepSeek R1 0528 | novita | 35.7M | 10.0M | 45.7M |
| 127 | Mixtral 8x22B | together | 20.8M | 20.8M | 41.7M |
| 128 | Qwen 2.5 72B | together | 20.8M | 20.8M | 41.7M |
| 129 | Cogito v2.1 671B | together | 20.0M | 20.0M | 40.0M |
| 130 | DeepSeek V3 | together | 20.0M | 20.0M | 40.0M |
| 131 | Gemini 3.6 Flash | 33.3M | 6.7M | 40.0M | |
| 132 | Gemini 3.7 Flash | 33.3M | 6.7M | 40.0M | |
| 133 | GPT-5.4 mini | openai | 33.3M | 5.6M | 38.9M |
| 134 | Kimi K2.6 | novita | 31.3M | 7.4M | 38.6M |
| 135 | gpt-3.5-turbo-1106 | openai | 25.0M | 12.5M | 37.5M |
| 136 | Grok Build 0.1 | xai | 25.0M | 12.5M | 37.5M |
| 137 | GLM-5 | together | 25.0M | 7.8M | 32.8M |
| 138 | GLM-5 | novita | 25.0M | 7.8M | 32.8M |
| 139 | Kimi K2.7 Code | together | 26.3M | 6.3M | 32.6M |
| 140 | Kimi K2.7 Code | novita | 26.3M | 6.3M | 32.6M |
| 141 | Qwen3 VL 235B A22B Thinking | novita | 25.5M | 6.3M | 31.8M |
| 142 | Claude Haiku 4.5 | anthropic | 25.0M | 5.0M | 30.0M |
| 143 | Grok 4.20 0309 Non-Reasoning | xai | 20.0M | 10.0M | 30.0M |
| 144 | Grok 4.20 0309 Reasoning | xai | 20.0M | 10.0M | 30.0M |
| 145 | Grok 4.20 Multi-Agent 0309 | xai | 20.0M | 10.0M | 30.0M |
| 146 | Grok 4.3 | xai | 20.0M | 10.0M | 30.0M |
| 147 | gpt-3.5-turbo-instruct | openai | 16.7M | 12.5M | 29.2M |
| 148 | o3-mini | openai | 22.7M | 5.7M | 28.4M |
| 149 | o4-mini | openai | 22.7M | 5.7M | 28.4M |
| 150 | Qwen3.7-Max | alibaba | 20.0M | 6.7M | 26.7M |
| 151 | Qwen3.7-Max | novita | 20.0M | 6.7M | 26.7M |
| 152 | Kimi K2.6 | together | 20.8M | 5.6M | 26.4M |
| 153 | DeepSeek V4 Pro 0813 | novita | 18.9M | 6.3M | 25.3M |
| 154 | GLM-5.1 | novita | 18.1M | 5.7M | 23.8M |
| 155 | GLM 5.3 | novita | 17.9M | 5.7M | 23.5M |
| 156 | GLM-5.1 | together | 17.9M | 5.7M | 23.5M |
| 157 | GLM-5.2 | together | 17.9M | 5.7M | 23.5M |
| 158 | GLM-5.2 | zai | 17.9M | 5.7M | 23.5M |
| 159 | GLM-5.2 | novita | 17.9M | 5.7M | 23.5M |
| 160 | GLM-5.3 | zai | 17.9M | 5.7M | 23.5M |
| 161 | DeepSeek V4 Pro | novita | 15.6M | 7.8M | 23.4M |
| 162 | Gemini 2.5 Pro | 20.0M | 2.5M | 22.5M | |
| 163 | GPT-5 | openai | 20.0M | 2.5M | 22.5M |
| 164 | GPT-5.1 | openai | 20.0M | 2.5M | 22.5M |
| 165 | DeepSeek V4 Pro | together | 14.4M | 7.2M | 21.6M |
| 166 | Mistral Medium 3.5 | mistral | 16.7M | 3.3M | 20.0M |
| 167 | Gemini 3.5 Flash | 16.7M | 2.8M | 19.4M | |
| 168 | Magistral Medium | mistral | 12.5M | 5.0M | 17.5M |
| 169 | Grok 4.5 | xai | 12.5M | 4.2M | 16.7M |
| 170 | Grok 4.6 | xai | 12.5M | 4.2M | 16.7M |
| 171 | Mixtral 8x22B | mistral | 12.5M | 4.2M | 16.7M |
| 172 | Pixtral Large | mistral | 12.5M | 4.2M | 16.7M |
| 173 | Qwen3.8 2.4T A95B | novita | 12.5M | 4.2M | 16.7M |
| 174 | Qwen3.8 Max | novita | 12.5M | 4.2M | 16.7M |
| 175 | GPT-5.2 | openai | 14.3M | 1.8M | 16.1M |
| 176 | GPT-4.1 | openai | 12.5M | 3.1M | 15.6M |
| 177 | o3 | openai | 12.5M | 3.1M | 15.6M |
| 178 | Sonar Deep Research | perplexity | 12.5M | 3.1M | 15.6M |
| 179 | Sonar Reasoning Pro | perplexity | 12.5M | 3.1M | 15.6M |
| 180 | Claude Sonnet 5 | anthropic | 12.5M | 2.5M | 15.0M |
| 181 | Gemini 3 Pro | 12.5M | 2.1M | 14.6M | |
| 182 | Gemini 3.1 Pro | 12.5M | 2.1M | 14.6M | |
| 183 | GPT-5.6 Terra | openai | 12.5M | 2.1M | 14.6M |
| 184 | Llama 3.1 405B | together | 7.1M | 7.1M | 14.3M |
| 185 | Command A | cohere | 10.0M | 2.5M | 12.5M |
| 186 | Command R+ 08-2024 | cohere | 10.0M | 2.5M | 12.5M |
| 187 | GPT-4o | openai | 10.0M | 2.5M | 12.5M |
| 188 | DeepSeek R1 | together | 8.3M | 3.6M | 11.9M |
| 189 | Gemini 2.5 Pro (>200k tokens) | 10.0M | 1.7M | 11.7M | |
| 190 | GPT-5.4 | openai | 10.0M | 1.7M | 11.7M |
| 191 | Kimi K3 | telnyx | 9.3M | 1.9M | 11.1M |
| 192 | Mistral Large | together | 8.3M | 2.8M | 11.1M |
| 193 | Claude Sonnet 4 | anthropic | 8.3M | 1.7M | 10.0M |
| 194 | Claude Sonnet 4.5 | anthropic | 8.3M | 1.7M | 10.0M |
| 195 | Claude Sonnet 4.6 | anthropic | 8.3M | 1.7M | 10.0M |
| 196 | Kimi K3 | moonshot | 8.3M | 1.7M | 10.0M |
| 197 | Kimi K3 | novita | 8.3M | 1.7M | 10.0M |
| 198 | Sonar Pro | perplexity | 8.3M | 1.7M | 10.0M |
| 199 | GPT-5.6 Sol | openai | 6.3M | 1.3M | 7.5M |
| 200 | gpt-4o-2024-05-13 | openai | 5.0M | 1.7M | 6.7M |
| 201 | Claude Opus 4.5 | anthropic | 5.0M | 1.0M | 6.0M |
| 202 | Claude Opus 4.6 | anthropic | 5.0M | 1.0M | 6.0M |
| 203 | Claude Opus 4.7 | anthropic | 5.0M | 1.0M | 6.0M |
| 204 | Claude Opus 4.8 | anthropic | 5.0M | 1.0M | 6.0M |
| 205 | Claude Opus 5 | anthropic | 5.0M | 1.0M | 6.0M |
| 206 | GPT-5.5 | openai | 5.0M | 833.3K | 5.8M |
| 207 | gpt-4-turbo-2024-04-09 | openai | 2.5M | 833.3K | 3.3M |
| 208 | Claude Fable 5 | anthropic | 2.5M | 500.0K | 3.0M |
| 209 | Claude Mythos 5 | anthropic | 2.5M | 500.0K | 3.0M |
| 210 | GPT-5.5 Cyber | openai | 2.0M | 333.3K | 2.3M |
| 211 | GPT-5.6 Cyber | openai | 2.0M | 333.3K | 2.3M |
| 212 | o1 | openai | 1.7M | 416.7K | 2.1M |
| 213 | Claude Opus 4 | anthropic | 1.7M | 333.3K | 2.0M |
| 214 | Claude Opus 4.1 | anthropic | 1.7M | 333.3K | 2.0M |
| 215 | GPT-5 Pro | openai | 1.7M | 208.3K | 1.9M |
| 216 | o3-pro | openai | 1.3M | 312.5K | 1.6M |
| 217 | GPT-5.2 Pro | openai | 1.2M | 148.8K | 1.3M |
| 218 | gpt-4-0613 | openai | 833.3K | 416.7K | 1.3M |
| 219 | GPT-5.4 Pro | openai | 833.3K | 138.9K | 972.2K |
| 220 | GPT-5.5 Pro | openai | 833.3K | 138.9K | 972.2K |
How does this calculator work?
- Enter your monthly budget — the total dollar amount you want to spend on AI API calls per month.
- Optionally select a focus model — highlights that model in the table and shows a detailed breakdown.
- Choose a budget split — decide how to allocate between input and output tokens (50/50, 80/20, or 20/80).
- Compare the table — models are ranked from most tokens to least, so the best-value models appear first.
Methodology
Token allowance is calculated by inverting the standard cost formula:
input_tokens = (budget × split_ratio / input_rate_per_M) × 1,000,000
output_tokens = (budget × (1 - split_ratio) / output_rate_per_M) × 1,000,000
The budget split controls what fraction of your monthly spend goes toward input vs output tokens. For most chat use cases, 50/50 is a reasonable default. If you send long prompts with short replies, use 80/20. If you request long-form content generation, use 20/80.
All rates come from our daily-updated pricing database. Models with $0 rates (free tiers) are excluded from ranking.
Does the budget only work at huge volume? If your workload is steady enough to keep GPUs busy, compare API spend with the GPU break-even guide, then test RunPod pricing against your real throughput.
The RunPod route may show a $5 referral credit after the first $10 added, but verify current terms and normal hourly pricing before treating it as part of your budget.
Affiliate disclosure: this link may earn us a commission at no extra cost to you. It does not affect token allowance rankings.