---
title: Cerebras
slug: cerebras
url: "https://toolweight.com/options/cerebras"
homepage: "https://cerebras.ai"
categories: llm-apis
last_verified: 2026-01-15
license: CC-BY-4.0
---

# Cerebras

> Wafer-scale inference, the fastest tokens per second available

Cerebras serves open-weight models from wafer-scale engines and consistently posts the highest single-stream throughput measured in this category, often several thousand tokens per second on models where GPU hosts manage a few hundred. Like Groq it is a latency instrument rather than a general platform: a short model list, and an architecture decision you make deliberately.

## Identity

|  |  |
| --- | --- |
| Name | Cerebras |
| Company | Cerebras |
| One-liner | Wafer-scale inference, the fastest tokens per second available |
| Site | https://cerebras.ai |
| Docs | https://inference-docs.cerebras.ai |
| Founded | 2016 |
| Funding | Late-stage private |
| Open source | No |
| Brand | https://toolweight.com/vendors/Cerebras |
| Compared in | 1 |

## Where it is compared

### [LLM APIs](https://toolweight.com/compare/llm-apis)
Ranked **#17 of 18** on default weights.
| Field | Value | Confidence | Verified | Source | Note |
| --- | --- | --- | --- | --- | --- |
| Flagship model | Curated open-weight models (Qwen3, GLM, Llama, GPT-OSS) | Vendor-claimed | 2026-01-15 | https://inference-docs.cerebras.ai/models/overview | - |
| Context window | - | Unknown | - | - | Per-model. |
| Max output | - | Unknown | - | - | Per-model. |
| Image input | ○ | Inferred | 2026-01-15 | - | - |
| Audio in/out | ○ | Inferred | 2026-01-15 | - | - |
| Open weights | ● | Inferred | 2026-01-15 | - | - |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://inference-docs.cerebras.ai/resources/openai | - |
| TTFT p50 | 200 ms | Community-reported | 2026-01-15 | - | Community-measured on short prompts; not measured by toolweight. |
| Output tok/s | 2,000 tok/s | Community-reported | 2026-01-15 | - | Order-of-magnitude figure, third-party benchmarks routinely report 1,500-3,000 tok/s on large MoE models where GPU hosts manage a few hundred. |
| $/M input | - | Unknown | - | - | Per-model. |
| $/M output | - | Unknown | - | - | - |
| $/M cache read | - | Unknown | - | - | - |
| Batch discount | - | Unknown | - | - | - |
| Tool use | ◐ | Inferred | 2026-01-15 | - | - |
| Schema output | ◐ | Inferred | 2026-01-15 | - | - |
| Effort control | ○ | Inferred | 2026-01-15 | - | - |
| Computer use | ○ | Inferred | 2026-01-15 | - | - |
| MCP support | ○ | Inferred | 2026-01-15 | - | - |
| Cache TTL | - | Unknown | - | - | - |
| Continuity policy | Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite. | Community-reported | 2026-01-15 | - | - |
| Notice period | - | Unknown | - | - | - |
| Pinnable versions | ◐ | Inferred | 2026-01-15 | - | - |
| Zero retention | - | Unknown | - | - | No contractual zero-retention programme was located, and this column's bar includes abuse-monitoring logs. Regraded from an unevidenced yes. |
| Trains on your data | ○ | Inferred | 2026-01-15 | - | - |
| AU region | ○ | Inferred | 2026-01-15 | - | - |
| API since | 2024 | Inferred | - | - | - |
| Positioning | Wafer-scale inference, the highest tokens per second available | Inferred | - | - | - |

**Verdict.** Consistently the fastest single-stream generation you can buy, by a margin that is qualitative rather than incremental, reasoning models that take half a minute elsewhere return in a couple of seconds. The catch is the same as Groq's: a short menu, no bring-your-own-weights, and no continuity guarantees at all.

## Alternatives

- [Qwen](https://toolweight.com/options/alibaba-qwen), Alibaba's model family, huge open-weight range, closed flagship
- [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock), Multi-vendor model access inside your existing AWS account
- [Anthropic](https://toolweight.com/options/anthropic), Claude models, built around long agentic runs and tool use
- [Cohere](https://toolweight.com/options/cohere), Enterprise-focused models built for RAG and private deployment
- [DeepSeek](https://toolweight.com/options/deepseek), Frontier-adjacent models at a small fraction of Western prices
- [Fireworks AI](https://toolweight.com/options/fireworks-ai), Fast open-weight inference with strong structured-output support
- [Google Gemini](https://toolweight.com/options/google-gemini), Gemini via AI Studio for prototyping or Vertex AI for production
- [Groq](https://toolweight.com/options/groq), Custom LPU silicon serving open-weight models at extreme speed
- [Meta Llama](https://toolweight.com/options/meta-llama), Open-weight Llama models, hosted almost everywhere but Meta
- [Mistral AI](https://toolweight.com/options/mistral-ai), European lab with an open-weight lineage and EU-resident hosting
- [Moonshot AI](https://toolweight.com/options/moonshot-ai), Kimi models, open-weight agentic performance at low cost
- [OpenAI](https://toolweight.com/options/openai), GPT models plus audio, images and embeddings on one bill

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/options/cerebras
- Machine-readable: https://toolweight.com/options/cerebras.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
