---
title: Groq
slug: groq
url: "https://toolweight.com/options/groq"
homepage: "https://groq.com"
categories: llm-apis
last_verified: 2026-01-15
license: CC-BY-4.0
---

# Groq

> Custom LPU silicon serving open-weight models at extreme speed

Groq runs a curated set of open-weight models on its own language processing units, delivering token rates that general-purpose GPU hosts do not approach. The API is OpenAI-compatible and cheap. The constraint is the menu: you get the models Groq has chosen to deploy, models rotate off with limited notice, and there is no path to bring your own weights.

## Identity

|  |  |
| --- | --- |
| Name | Groq |
| Company | Groq |
| One-liner | Custom LPU silicon serving open-weight models at extreme speed |
| Site | https://groq.com |
| Docs | https://console.groq.com/docs |
| Founded | 2016 |
| Funding | Late-stage private |
| Open source | No |
| Brand | https://toolweight.com/vendors/Groq |
| Compared in | 1 |

## Where it is compared

### [LLM APIs](https://toolweight.com/compare/llm-apis)
Ranked **#16 of 18** on default weights.
| Field | Value | Confidence | Verified | Source | Note |
| --- | --- | --- | --- | --- | --- |
| Flagship model | Curated open-weight models (Kimi K2, Llama 4, Qwen3, GPT-OSS) | Vendor-claimed | 2026-01-15 | https://console.groq.com/docs/models | The menu is short and rotates; check the models page before designing around one. |
| Context window | - | Unknown | - | - | Per-model. |
| Max output | - | Unknown | - | - | Per-model. |
| Image input | ◐ | Inferred | 2026-01-15 | - | - |
| Audio in/out | ◐ | Inferred | 2026-01-15 | - | Whisper-class transcription is served fast; no native speech output. |
| Open weights | ● | Inferred | 2026-01-15 | - | - |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://console.groq.com/docs/openai | - |
| TTFT p50 | 250 ms | Community-reported | 2026-01-15 | - | Community-measured on short prompts from third-party benchmarks; varies by model and region. Not measured by toolweight. |
| Output tok/s | 400 tok/s | Community-reported | 2026-01-15 | - | Order-of-magnitude figure for mid-sized models, roughly 5-10x a general-purpose GPU host. Smaller models run considerably faster. |
| $/M input | - | Unknown | - | - | Per-model and consistently below GPU hosts for the same weights. |
| $/M output | - | Unknown | - | - | - |
| $/M cache read | - | Unknown | - | - | - |
| Batch discount | - | Unknown | - | - | A batch API is offered at a discount, which is unusual among the speed-focused hosts, but the percentage is disputed. The 50% previously carried here is the category norm rather than a figure confirmed for Groq, and a materially lower rate has been reported. Since this is the only pricing-family number any router or host on this page carries, it is blank rather than assumed: verify against the console before modelling a saving on it. |
| Tool use | ◐ | Inferred | 2026-01-15 | - | - |
| Schema output | ◐ | Inferred | 2026-01-15 | - | - |
| Effort control | ○ | Inferred | 2026-01-15 | - | - |
| Computer use | ○ | Inferred | 2026-01-15 | - | - |
| MCP support | ○ | Inferred | 2026-01-15 | - | - |
| Cache TTL | - | Unknown | - | - | - |
| Continuity policy | The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun. | Community-reported | 2026-01-15 | - | - |
| Notice period | - | Unknown | - | - | Deprecations are announced but the window is short and not published as a guaranteed minimum. |
| Pinnable versions | ◐ | Inferred | 2026-01-15 | - | You can name a specific model, but it only exists while Groq chooses to host it. |
| Zero retention | - | Unknown | - | - | No contractual zero-retention programme was located, and this column's bar includes abuse-monitoring logs. Regraded from an unevidenced yes. |
| Trains on your data | ○ | Inferred | 2026-01-15 | - | - |
| AU region | ○ | Inferred | 2026-01-15 | - | - |
| API since | 2024 | Inferred | - | - | - |
| Positioning | Custom LPU silicon for open-weight models at very low latency | Inferred | - | - | - |

**Verdict.** When the response appearing instantly is the product, Groq changes what you can build, voice, live search and inline completion feel different at these token rates. It is not a general-purpose platform: the menu is short, models rotate off with little warning, and there is no path to bring your own weights.

## Alternatives

- [Qwen](https://toolweight.com/options/alibaba-qwen), Alibaba's model family, huge open-weight range, closed flagship
- [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock), Multi-vendor model access inside your existing AWS account
- [Anthropic](https://toolweight.com/options/anthropic), Claude models, built around long agentic runs and tool use
- [Cerebras](https://toolweight.com/options/cerebras), Wafer-scale inference, the fastest tokens per second available
- [Cohere](https://toolweight.com/options/cohere), Enterprise-focused models built for RAG and private deployment
- [DeepSeek](https://toolweight.com/options/deepseek), Frontier-adjacent models at a small fraction of Western prices
- [Fireworks AI](https://toolweight.com/options/fireworks-ai), Fast open-weight inference with strong structured-output support
- [Google Gemini](https://toolweight.com/options/google-gemini), Gemini via AI Studio for prototyping or Vertex AI for production
- [Meta Llama](https://toolweight.com/options/meta-llama), Open-weight Llama models, hosted almost everywhere but Meta
- [Mistral AI](https://toolweight.com/options/mistral-ai), European lab with an open-weight lineage and EU-resident hosting
- [Moonshot AI](https://toolweight.com/options/moonshot-ai), Kimi models, open-weight agentic performance at low cost
- [OpenAI](https://toolweight.com/options/openai), GPT models plus audio, images and embeddings on one bill

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/options/groq
- Machine-readable: https://toolweight.com/options/groq.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
