LLM API pricing, compared
Pick a scale or set your usage. Each tool shows its cheapest plan for it, cheapest first, with the arithmetic and the vendor's pricing page. Price is one input — which one fits your situation matters as much.
Mistral Large 3 ($0.50 / $1.50) is still sold but superseded by Large 4. OCR and voice models are priced per page or per minute and aren't modeled here.
Standard-tier prices for models also sold by their own labs; many more open models are listed. Cached input is billed lower.
Prices from Groq's production model list (groq.com/pricing no longer shows a price table). Llama 3.1 8B and Llama 3.3 70B are listed as contact sales.
International (BytePlus) prices for the lowest input-length tier. Flex (off-peak) online inference and batch inference cost about half. Mainland prices (Volcano Engine Ark) are in yuan.
Cache-miss input prices. Cache hits cost a fraction of a cent (Pro $0.0036, Flash $0.0028 per million) and cache writes are free for now. The Batch API halves Pro and Flash prices. Mainland prices are in yuan, and a separate Token Plan subscription covers coding tools.
Base input/output prices; prompt-cache reads are billed much lower and cache writes higher. US-only inference is 1.1x, and fast mode for Opus 5.5 is 2x. Batch processing is discounted.
International prices for the lowest input-length tier. ERNIE 5.0, which adds image input and a thinking mode, costs $0.84 / $3.38.
Standard-tier prices for short context (the page doesn't state the cutoff for GPT-6 models; older models switch at 272K input tokens); long-context requests cost more (Astra $20 / $75, Sol $4 / $15, Luna $0.20 / $0.75 per million). Batch is half price; cached input is billed lower.
Cached input is billed lower (Hy4 preview $0.042, Hy3 $0.033 per million). Hy4 is still a preview release.
International (Singapore) list prices for the lowest input-length tier. Qwen3.7-Plus is marked "limited-time 20% off" on the page; the list price is shown. Batch calls are half price.
Serverless standard prices for the models also sold by their own labs. Models not listed are priced by size, from $0.10 per million tokens under 4B parameters.
International (Z.ai) prices, flat across input lengths. GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free. Mainland prices (bigmodel.cn) are in yuan and tiered by length for some models.
Serverless prices for the models also sold by their own labs; many more open models are listed on the page. Cached input is billed lower where shown, and the Batch API is cheaper.
A rate-limited free trial (about 1M tokens a day) and a one-time $5 credit are available. Qwen 3.8 27B is the other shared model, at $0.99 / $1.49.
Peak-hour prices with cache-miss input (peak is 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday). All other hours are half price (v4-pro $0.66 / $1.98, flash $0.15 / $0.60), and cache-hit input is far cheaper, so most workloads pay less than shown.
Cache reads cost $0.06 per million. The small MiniMax-M3.1-Flash-Preview is sold only through subscription plans, not pay-as-you-go. Mainland prices (platform.minimaxi.com) are in yuan.
Paid-tier Standard text prices. Gemini 3.8 Flash is $0.75 / $3.75 through December 31, 2026 and rises to $1.50 / $7.50 on January 1, 2027. Gemini 3.1 Pro Preview costs $4 / $18 for prompts over 200K tokens. Batch and Flex are 50% off. A free tier exists but may use prompts to improve Google's products.
Prices for prompts under 200K tokens; at 200K and above input and output prices double. Web and X search tools are billed separately as tool calls.
Cache-miss input prices. Cache hits are far cheaper (K3 $0.30, K2.6 $0.16, K2.7 Code $0.19 per million). A HighSpeed variant of K2.7 Code costs twice as much. Mainland prices (platform.moonshot.cn) are in yuan.
Standard-tier prices for muse-spark-1.3 (also 1.2 and 1.1); cached input is $0.15 per million and there is no long-context premium. The contributor tier is left out because it trades your data for the lower price.
Read from each vendor's pricing page (oldest check 2026-10-03). The model is plan fee + usage beyond what's included × the vendor's rate; volume discounts, annual billing, taxes and regional prices are left out — see each tool's notes. Always confirm on the vendor's page before you commit.
Cost as you grow
Each tool's cheapest usable plan as the product grows.
At three scales
| Tool | Side projectinput tokens 5 million tokens, output tokens 1 million tokens | Growinginput tokens 50 million tokens, output tokens 10 million tokens | Scalinginput tokens 1 billion tokens, output tokens 200 million tokens |
|---|---|---|---|
| $0.60Ministral 3 (3B) | $6.00Ministral 3 (3B) | $120Ministral 3 (3B) | |
| DeepInfra | $0.63DeepSeek V4 Flash | $6.30DeepSeek V4 Flash | $126DeepSeek V4 Flash |
| $0.68GPT OSS 20B | $6.75GPT OSS 20B | $135GPT OSS 20B | |
| Doubao API | $0.90Seed 2.0 Mini | $9.00Seed 2.0 Mini | $180Seed 2.0 Mini |
| Xiaomi MiMo API | $0.98MiMo-V2.6-Flash | $9.80MiMo-V2.6-Flash | $196MiMo-V2.6-Flash |
| $1.00Claude Haiku 5.5 | $10Claude Haiku 5.5 | $200Claude Haiku 5.5 | |
| ERNIE API | $1.00ERNIE 4.5 Turbo | $10ERNIE 4.5 Turbo | $200ERNIE 4.5 Turbo |
| $1.00GPT-6 Luna | $10GPT-6 Luna | $200GPT-6 Luna | |
| Hunyuan API | $1.19Hy3 | $12Hy3 | $238Hy3 |
| Qwen API | $1.22Qwen3.8-Flash | $12Qwen3.8-Flash | $244Qwen3.8-Flash |
| $1.25GLM 5.3 Flash | $13GLM 5.3 Flash | $250GLM 5.3 Flash | |
| GLM API | $1.25GLM-5.3-Flash | $13GLM-5.3-Flash | $250GLM-5.3-Flash |
| $1.35gpt-oss-120B | $14gpt-oss-120B | $270gpt-oss-120B | |
| Cerebras Inference API | $2.50gpt-oss-120B | $25gpt-oss-120B | $500gpt-oss-120B |
| $2.70deepseek-flash | $27deepseek-flash | $540deepseek-flash | |
| MiniMax | $2.70MiniMax-M3 | $27MiniMax-M3 | $540MiniMax-M3 |
| $2.75Gemini 3.1 Flash-Lite | $28Gemini 3.1 Flash-Lite | $550Gemini 3.1 Flash-Lite | |
| $8.75Grok 4.3 | $88Grok 4.3 | $1,750Grok 4.3 | |
| Kimi API | $8.75Kimi K2.6 | $88Kimi K2.6 | $1,750Kimi K2.6 |
| Meta Model API | $11Muse Spark | $105Muse Spark | $2,100Muse Spark |
Every plan
What each LLM API plan costs, includes and caps, as modelled here.
Mistral AIPay per token, from $0.10 per million tokens on Ministral 3 (3B) to $1.50 / $7.50 per million input/output tokens on Mistral Medium 3.5.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Mistral Large 4 | None | $1.36 per million tokens | $4.18 per million tokens |
| Mistral Medium 3.5 | None | $1.5 per million tokens | $7.5 per million tokens |
| Mistral Small 4 | None | $0.15 per million tokens | $0.6 per million tokens |
| Ministral 3 (3B) | None | $0.1 per million tokens | $0.1 per million tokens |
Mistral Large 3 ($0.50 / $1.50) is still sold but superseded by Large 4. OCR and voice models are priced per page or per minute and aren't modeled here.
Checked 2026-10-08 · pricing page ↗ · Mistral AI pricing in full →DeepInfraPay per token for open-weight models, e.g. $0.09 / $0.18 per million input/output tokens on DeepSeek V4 Flash and $2.85 / $14.25 on Kimi K3.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Kimi K3 | None | $2.85 per million tokens | $14.25 per million tokens |
| Qwen3.8-Max | None | $1.65 per million tokens | $4.951 per million tokens |
| DeepSeek V4 Pro | None | $1.3 per million tokens | $2.6 per million tokens |
| DeepSeek V4 Flash | None | $0.09 per million tokens | $0.18 per million tokens |
| Nemotron 3 Ultra | None | $0.5 per million tokens | $2.2 per million tokens |
| Muse Glimmer 30B | None | $0.3 per million tokens | $1.2 per million tokens |
Standard-tier prices for models also sold by their own labs; many more open models are listed. Cached input is billed lower.
Checked 2026-10-07 · pricing page ↗ · DeepInfra pricing in full →
GroqPay per token on production models: $0.15 / $0.60 per million input/output tokens for GPT OSS 120B and $0.075 / $0.30 for GPT OSS 20B.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| GPT OSS 120B | None | $0.15 per million tokens | $0.6 per million tokens |
| GPT OSS 20B | None | $0.075 per million tokens | $0.3 per million tokens |
Prices from Groq's production model list (groq.com/pricing no longer shows a price table). Llama 3.1 8B and Llama 3.3 70B are listed as contact sales.
Checked 2026-10-03 · pricing page ↗ · Groq pricing in full →Doubao APIPay per token on BytePlus ModelArk, from $0.10 / $0.40 per million input/output tokens on Seed 2.0 Mini to $0.50 / $3 on Seed 2.0 Pro.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Seed 2.0 Pro | None | $0.5 per million tokens | $3 per million tokens |
| Seed 2.1 Turbo | None | $0.5 per million tokens | $2.5 per million tokens |
| Seed 2.0 Lite | None | $0.25 per million tokens | $2 per million tokens |
| Seed 2.0 Mini | None | $0.1 per million tokens | $0.4 per million tokens |
International (BytePlus) prices for the lowest input-length tier. Flex (off-peak) online inference and batch inference cost about half. Mainland prices (Volcano Engine Ark) are in yuan.
Checked 2026-10-07 · pricing page ↗ · Doubao API pricing in full →Xiaomi MiMo APIPay per token from a prepaid balance, from $0.14 / $0.28 per million input/output tokens on MiMo-V2.6-Flash to $0.435 / $0.87 on MiMo-V2.6-Pro.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| MiMo-V2.6-Pro | None | $0.435 per million tokens | $0.87 per million tokens |
| MiMo-V2.6-Flash | None | $0.14 per million tokens | $0.28 per million tokens |
Cache-miss input prices. Cache hits cost a fraction of a cent (Pro $0.0036, Flash $0.0028 per million) and cache writes are free for now. The Batch API halves Pro and Flash prices. Mainland prices are in yuan, and a separate Token Plan subscription covers coding tools.
Checked 2026-10-08 · pricing page ↗ · Xiaomi MiMo API pricing in full →
ClaudePay per token with no monthly fee, from $0.10 / $0.50 per million input/output tokens on Haiku 5.5 to $10 / $50 on Fable 5.1.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Claude Fable 5.1 | None | $10 per million tokens | $50 per million tokens |
| Claude Opus 5.5 | None | $4 per million tokens | $20 per million tokens |
| Claude Sonnet 5.5 | None | $2 per million tokens | $10 per million tokens |
| Claude Haiku 5.5 | None | $0.1 per million tokens | $0.5 per million tokens |
Base input/output prices; prompt-cache reads are billed much lower and cache writes higher. US-only inference is 1.1x, and fast mode for Opus 5.5 is 2x. Batch processing is discounted.
Checked 2026-10-08 · pricing page ↗ · Claude pricing in full →ERNIE APIPay per token on Qianfan, from $0.11 / $0.45 per million input/output tokens on ERNIE 4.5 Turbo to $0.56 / $2.53 on ERNIE 5.1.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| ERNIE 5.1 | None | $0.56 per million tokens | $2.53 per million tokens |
| ERNIE 4.5 Turbo | None | $0.11 per million tokens | $0.45 per million tokens |
International prices for the lowest input-length tier. ERNIE 5.0, which adds image input and a thinking mode, costs $0.84 / $3.38.
Checked 2026-10-07 · pricing page ↗ · ERNIE API pricing in full →
OpenAI APIPay per token, from $0.10 / $0.50 per million input/output tokens on GPT-6 Luna to $10 / $50 on GPT-6 Astra.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| GPT-6 Astra | None | $10 per million tokens | $50 per million tokens |
| GPT-6.1 Sol | None | $2 per million tokens | $10 per million tokens |
| GPT-6 Luna | None | $0.1 per million tokens | $0.5 per million tokens |
Standard-tier prices for short context (the page doesn't state the cutoff for GPT-6 models; older models switch at 272K input tokens); long-context requests cost more (Astra $20 / $75, Sol $4 / $15, Luna $0.20 / $0.75 per million). Batch is half price; cached input is billed lower.
Checked 2026-10-03 · pricing page ↗ · OpenAI API pricing in full →Hunyuan APIPay per token, from $0.132 / $0.528 per million input/output tokens on Hy3 to $0.834 / $2.501 on Hy4 preview.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Hy4 preview | None | $0.834 per million tokens | $2.501 per million tokens |
| Hy3 | None | $0.132 per million tokens | $0.528 per million tokens |
Cached input is billed lower (Hy4 preview $0.042, Hy3 $0.033 per million). Hy4 is still a preview release.
Checked 2026-10-07 · pricing page ↗ · Hunyuan API pricing in full →Qwen APIPay per token, from $0.15 / $0.47 per million input/output tokens on Qwen3.8-Flash to $2 / $6 on Qwen3.8-Max (international region).
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Qwen3.8-Max | None | $2 per million tokens | $6 per million tokens |
| Qwen3.7-Plus | None | $0.4 per million tokens | $1.6 per million tokens |
| Qwen3.8-Flash | None | $0.15 per million tokens | $0.47 per million tokens |
International (Singapore) list prices for the lowest input-length tier. Qwen3.7-Plus is marked "limited-time 20% off" on the page; the list price is shown. Batch calls are half price.
Checked 2026-10-07 · pricing page ↗ · Qwen API pricing in full →
Fireworks AIPay per token for other labs' open-weight models, e.g. $0.15 / $0.60 per million input/output tokens on gpt-oss-120B and $3 / $15 on Kimi K3.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Kimi K3 | None | $3 per million tokens | $15 per million tokens |
| Qwen 3.8 Max | None | $2 per million tokens | $6 per million tokens |
| GLM 5.3 | None | $1.4 per million tokens | $4.4 per million tokens |
| MiniMax M3 | None | $0.3 per million tokens | $1.2 per million tokens |
| DeepSeek V4.1 Flash | None | $0.3 per million tokens | $1.2 per million tokens |
| GLM 5.3 Flash | None | $0.15 per million tokens | $0.5 per million tokens |
| gpt-oss-120B | None | $0.15 per million tokens | $0.6 per million tokens |
Serverless standard prices for the models also sold by their own labs. Models not listed are priced by size, from $0.10 per million tokens under 4B parameters.
Checked 2026-10-07 · pricing page ↗ · Fireworks AI pricing in full →GLM APIPay per token, from $0.15 / $0.50 per million input/output tokens on GLM-5.3-Flash to $1.40 / $4.40 on GLM-5.3, with some older Flash models free.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| GLM-5.3 | None | $1.4 per million tokens | $4.4 per million tokens |
| GLM-5.3-FlashX | None | $0.37 per million tokens | $1.25 per million tokens |
| GLM-5.3-Flash | None | $0.15 per million tokens | $0.5 per million tokens |
International (Z.ai) prices, flat across input lengths. GLM-4.7-Flash, GLM-4.5-Flash and GLM-4.6V-Flash are free. Mainland prices (bigmodel.cn) are in yuan and tiered by length for some models.
Checked 2026-10-07 · pricing page ↗ · GLM API pricing in full →
Together AIPay per token for other labs' open-weight models, e.g. $0.15 / $0.60 per million input/output tokens on gpt-oss-120B and $1.40 / $4.40 on GLM-5.3.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Kimi K3 | None | $2.7 per million tokens | $13.5 per million tokens |
| GLM-5.3 | None | $1.4 per million tokens | $4.4 per million tokens |
| DeepSeek V4 Pro | None | $1.32 per million tokens | $3.96 per million tokens |
| DeepSeek V4.1 Flash | None | $0.3 per million tokens | $1.2 per million tokens |
| gpt-oss-120B | None | $0.15 per million tokens | $0.6 per million tokens |
| Muse Glimmer 30B | None | $0.35 per million tokens | $1.5 per million tokens |
Serverless prices for the models also sold by their own labs; many more open models are listed on the page. Cached input is billed lower where shown, and the Batch API is cheaper.
Checked 2026-10-07 · pricing page ↗ · Together AI pricing in full →Cerebras Inference APIPay per token on very fast inference: $0.35 / $0.75 per million input/output tokens on gpt-oss-120B.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| gpt-oss-120B | None | $0.35 per million tokens | $0.75 per million tokens |
A rate-limited free trial (about 1M tokens a day) and a one-time $5 credit are available. Qwen 3.8 27B is the other shared model, at $0.99 / $1.49.
Checked 2026-10-07 · pricing page ↗ · Cerebras Inference API pricing in full →
DeepSeekPay per token from a prepaid balance, with off-peak hours at half price and cache-hit input far cheaper.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| deepseek-v4-pro | None | $1.32 per million tokens | $3.96 per million tokens |
| deepseek-flash | None | $0.3 per million tokens | $1.2 per million tokens |
Peak-hour prices with cache-miss input (peak is 01:00–04:00 and 06:00–10:00 UTC, Monday–Friday). All other hours are half price (v4-pro $0.66 / $1.98, flash $0.15 / $0.60), and cache-hit input is far cheaper, so most workloads pay less than shown.
Checked 2026-10-03 · pricing page ↗ · DeepSeek pricing in full →MiniMaxPay per token: $0.30 / $1.20 per million input/output tokens on MiniMax-M3 for requests up to 512K input tokens.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| MiniMax-M3 | None | $0.3 per million tokens | $1.2 per million tokens |
Cache reads cost $0.06 per million. The small MiniMax-M3.1-Flash-Preview is sold only through subscription plans, not pay-as-you-go. Mainland prices (platform.minimaxi.com) are in yuan.
Checked 2026-10-07 · pricing page ↗ · MiniMax pricing in full →
Gemini APIPay per token on the paid tier, from $0.25 / $1.50 per million input/output tokens on Gemini 3.1 Flash-Lite to $2 / $12 on 3.1 Pro Preview; a free tier covers some models.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Gemini 3.1 Pro Preview | None | $2 per million tokens | $12 per million tokens |
| Gemini 3.8 Flash | None | $0.75 per million tokens | $3.75 per million tokens |
| Gemini 3.5 Flash-Lite | None | $0.3 per million tokens | $2.5 per million tokens |
| Gemini 3.1 Flash-Lite | None | $0.25 per million tokens | $1.5 per million tokens |
Paid-tier Standard text prices. Gemini 3.8 Flash is $0.75 / $3.75 through December 31, 2026 and rises to $1.50 / $7.50 on January 1, 2027. Gemini 3.1 Pro Preview costs $4 / $18 for prompts over 200K tokens. Batch and Flex are 50% off. A free tier exists but may use prompts to improve Google's products.
Checked 2026-10-03 · pricing page ↗ · Gemini API pricing in full →
Grok APIPay per token: $2 / $6 per million input/output tokens on Grok 4.7 and $1.25 / $2.50 on Grok 4.3.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Grok 4.7 | None | $2 per million tokens | $6 per million tokens |
| Grok 4.3 | None | $1.25 per million tokens | $2.5 per million tokens |
Prices for prompts under 200K tokens; at 200K and above input and output prices double. Web and X search tools are billed separately as tool calls.
Checked 2026-10-03 · pricing page ↗ · Grok API pricing in full →Kimi APIPay per token from a prepaid balance, from $0.95 / $4 per million input/output tokens on Kimi K2.6 to $3 / $15 on Kimi K3.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Kimi K3 | None | $3 per million tokens | $15 per million tokens |
| Kimi K2.6 | None | $0.95 per million tokens | $4 per million tokens |
| Kimi K2.7 Code | None | $0.95 per million tokens | $4 per million tokens |
Cache-miss input prices. Cache hits are far cheaper (K3 $0.30, K2.6 $0.16, K2.7 Code $0.19 per million). A HighSpeed variant of K2.7 Code costs twice as much. Mainland prices (platform.moonshot.cn) are in yuan.
Checked 2026-10-07 · pricing page ↗ · Kimi API pricing in full →Meta Model APIPay per token, no minimums: $1.25 / $4.25 per million input/output tokens on Muse Spark, or far less on a tier that lets Meta train on your data.
| Plan | Monthly fee | Input tokens per month | Output tokens per month |
|---|---|---|---|
| Muse Spark | None | $1.25 per million tokens | $4.25 per million tokens |
Standard-tier prices for muse-spark-1.3 (also 1.2 and 1.1); cached input is $0.15 per million and there is no long-context premium. The contributor tier is left out because it trades your data for the lower price.
Checked 2026-10-07 · pricing page ↗ · Meta Model API pricing in full →