Cerebras

Wafer-scale inference — the fastest tokens per second available

Compared on toolweight in:LLM APIs
Related comparisons:Coding agents

Logo assets

LogoSVG
Download
Logo (dark)SVG
Download
WordmarkSVG
Download
Wordmark (dark)SVG
Download
MonochromeSVG

Flattened to currentColor — what toolweight renders in tables and nav.

Download

Cerebras brand colours

Extracted from the artwork and ranked by how much of the mark each covers — not from a brand book, so treat them as observed rather than official.

#F15A29
#000000

What Cerebras is

Cerebras serves open-weight models from wafer-scale engines and consistently posts the highest single-stream throughput measured in this category, often several thousand tokens per second on models where GPU hosts manage a few hundred. Like Groq it is a latency instrument rather than a general platform: a short model list, and an architecture decision you make deliberately.

Recent changes

  • 2025-08-05OpenAI releases open-weight gpt-oss models

    gpt-oss-120b and gpt-oss-20b shipped under Apache 2.0, OpenAI's first open weights since GPT-2. Both landed on Groq, Together, Fireworks and Cerebras within days, and reset expectations that the frontier labs would keep everything closed.

Full Cerebras changelog →

Logo artwork sourced from svgl, the MIT-licensed SVG logo library by pheralb. github.com/pheralb/svgl · MIT licence

Logos are trademarks of their respective owners. toolweight is not affiliated with, sponsored by, or endorsed by them. Marks are reproduced for identification in editorial comparison. Rights holders can request removal and we will honour it immediately.