Cerebras Inference API

Very fast inference on a small set of open-weight models, served on Cerebras's own wafer-scale chips through an OpenAI-compatible API.

Ask your AI about this, with this page as the source:ChatGPT ↗Claude ↗Perplexity ↗

Cerebras is a US chipmaker whose wafer-scale processors it also runs as a cloud. The Inference API serves open-weight models from other labs at very high output speeds — Cerebras lists around 3,000 tokens per second for OpenAI's gpt-oss-120b and around 1,850 for Qwen 3.8 27B, the two models on its self-serve shared tier. Many more model families, and fine-tuned weights, are available on Dedicated Inference by contract.

You call it with Cerebras's own SDKs (cerebras_cloud_sdk for Python, @cerebras/cerebras_cloud_sdk for Node) or with the OpenAI SDK by setting the base URL to https://api.cerebras.ai/v1; the Vercel AI SDK has a Cerebras provider as well. Pay-as-you-go usage is bought as prepaid credits, with optional auto-recharge.

Useful for a small team: speed, for features where users wait on long outputs or agents chain many calls; a Free Trial with $5 of credit for 30 days after you add a payment method; and once you buy credits, no hourly or daily token caps. Prompt caching is automatic, and cached tokens don't count toward the main uncached-token rate limit, so a high cache hit rate raises effective throughput. Shared models are served unpruned, and Cerebras commits not to change a model ID's architecture without notice.

The main limits: only two models on the self-serve tier, and anything else means talking to sales. Context is capped at 131k tokens on paid plans (about half that on the trial), cached input is billed at the full input rate, and the trial is meant for evaluation, not production.

Where it fits

Who uses it

No maker's live product on record yet, and 11 open-source projects that declare it in their code.

We haven't found a maker's live product that uses Cerebras Inference API yet — only the open-source projects below, which declare it in their code.

Open source: a project that declares Cerebras Inference API as a dependency in its public code — verifiable, but not necessarily a live product.

Alternatives to Cerebras Inference API

All alternatives by situation →

Questions makers ask about Cerebras Inference API

Can I use the OpenAI SDK?

Yes. Set the base URL to https://api.cerebras.ai/v1 and use a Cerebras API key. A few OpenAI features differ — for example, images must be sent as base64 data URIs rather than external URLs. source ↗

Is there a free tier?

There is a Free Trial, not a permanent free tier. New accounts get $5 of credits after adding a verified payment method; they expire after 30 days, and API access pauses when they run out until you buy credits. source ↗

Which models can I use without a contract?

The shared tier currently serves gpt-oss-120b and Qwen 3.8 27B, on both the Free Trial and pay-as-you-go. Other families — more Qwen, GPT-OSS, MiniMax and Gemma variants, and your own fine-tuned weights — run on Dedicated Inference. source ↗

What are the rate limits?

Limits apply per organization and per model, counted in requests and tokens per minute, with separate caps on uncached and total tokens. On pay-as-you-go, gpt-oss-120b allows 1,000 requests and 1M uncached tokens per minute, with no hourly or daily caps; the Free Trial is far lower, at 5 requests per minute and 1M tokens per day. source ↗

Does prompt caching lower the price?

No. Caching is automatic and reduces latency, but cached input tokens are billed at the normal input rate. Cached tokens don't count toward your uncached-token limit, and cache entries last at least 5 minutes. source ↗

Are the models compressed to make them faster?

Shared models are the original, unpruned versions. Cerebras stores some weights in lower precision and dequantizes them on the fly, while activations, attention and the KV cache stay at full precision. source ↗

Can I buy it through another platform?

Yes. Cerebras inference is also offered through AWS Marketplace, OpenRouter, Hugging Face and Vercel. source ↗

Is Cerebras Inference API free?

No permanent free tier; paid use starts at Pay per token. source ↗

Can AI coding agents work with Cerebras Inference API?

No llms.txt, official MCP server or CLI found yet.

Who uses Cerebras Inference API?

No maker's live product we track yet; 11 open-source projects declare it in their code. source ↗