Cerebras
Wafer-scale inference — the fastest tokens per second available
Logo assets
Cerebras brand colours
Extracted from the artwork and ranked by how much of the mark each covers — not from a brand book, so treat them as observed rather than official.
What Cerebras is
Cerebras serves open-weight models from wafer-scale engines and consistently posts the highest single-stream throughput measured in this category, often several thousand tokens per second on models where GPU hosts manage a few hundred. Like Groq it is a latency instrument rather than a general platform: a short model list, and an architecture decision you make deliberately.
Recent changes
- 2025-08-05OpenAI releases open-weight gpt-oss models
gpt-oss-120b and gpt-oss-20b shipped under Apache 2.0, OpenAI's first open weights since GPT-2. Both landed on Groq, Together, Fireworks and Cerebras within days, and reset expectations that the frontier labs would keep everything closed.