---
title: Frontier LLM APIs compared — price, context, tool use and deprecation policy
slug: llm-apis
url: "https://toolweight.dev/compare/llm-apis"
question: Which frontier LLM API should I build on?
tools: 18
fields: 27
last_verified: 2026-06-24
freshness: 27%
verify_cadence: 7d
license: CC-BY-4.0
---

# Frontier LLM API providers

> **Which frontier LLM API should I build on?**
>
> Anthropic, OpenAI and Google lead on capability; DeepSeek, Moonshot, Z.ai and Qwen undercut them by roughly an order of magnitude on price. Routers like OpenRouter, Groq and Cerebras trade first-party features for reach or speed. Pick on continuity and tool-use fidelity, not headline token price — price moves monthly, migrations do not.

## At a glance

|  |  |
| --- | --- |
| Canonical | https://toolweight.dev/compare/llm-apis |
| Tools compared | 18 |
| Fields compared | 27 |
| Last verified | 2026-06-24 (29d ago) |
| Verify cadence | 7d |
| Freshness | 27% (mean cell confidence) |
| Killer field | continuity_policy |
| Scoring | Weighted, re-rankable |
| Licence | CC-BY-4.0 — https://creativecommons.org/licenses/by/4.0/ |

## Ranking — default weights

| # | Tool | Score | Coverage | One-liner |
| --- | --- | --- | --- | --- |
| 1 | [Google Gemini](https://toolweight.dev/tools/google-gemini) | 79.9 | 83% | Gemini via AI Studio for prototyping or Vertex AI for production |
| 2 | [OpenAI](https://toolweight.dev/tools/openai) | 79.8 | 87% | GPT models plus audio, images and embeddings on one bill |
| 3 | [Amazon Bedrock](https://toolweight.dev/tools/amazon-bedrock) | 78.1 | 74% | Multi-vendor model access inside your existing AWS account |
| 4 | [Mistral AI](https://toolweight.dev/tools/mistral-ai) | 72.8 | 65% | European lab with an open-weight lineage and EU-resident hosting |
| 5 | [Qwen](https://toolweight.dev/tools/alibaba-qwen) | 67.4 | 74% | Alibaba's model family — huge open-weight range, closed flagship |
| 6 | [Meta Llama](https://toolweight.dev/tools/meta-llama) | 66.6 | 65% | Open-weight Llama models, hosted almost everywhere but Meta |
| 7 | [Self-hosted (vLLM)](https://toolweight.dev/tools/self-hosted-vllm) | 65.1 | 65% | Run open weights on your own GPUs behind an OpenAI-shaped API |
| 8 | [Together AI](https://toolweight.dev/tools/together-ai) | 64.2 | 61% | Serverless and dedicated hosting for open-weight models |
| 9 | [xAI](https://toolweight.dev/tools/xai) | 63.6 | 57% | Grok models with an OpenAI-shaped API and live X data access |
| 10 | [Fireworks AI](https://toolweight.dev/tools/fireworks-ai) | 63.2 | 57% | Fast open-weight inference with strong structured-output support |
| 11 | [Cohere](https://toolweight.dev/tools/cohere) | 62.9 | 74% | Enterprise-focused models built for RAG and private deployment |
| 12 | [Anthropic](https://toolweight.dev/tools/anthropic) | 62.7 | 91% | Claude models, built around long agentic runs and tool use |
| 13 | [OpenRouter](https://toolweight.dev/tools/openrouter) | 60.9 | 65% | One OpenAI-shaped key in front of hundreds of models |
| 14 | [Z.ai (GLM)](https://toolweight.dev/tools/zhipu-zai) | 59.9 | 74% | GLM models with an unusually cheap flat-rate coding plan |
| 15 | [Moonshot AI](https://toolweight.dev/tools/moonshot-ai) | 58.8 | 78% | Kimi models — open-weight agentic performance at low cost |
| 16 | [Groq](https://toolweight.dev/tools/groq) | 56.6 | 65% | Custom LPU silicon serving open-weight models at extreme speed |
| 17 | [Cerebras](https://toolweight.dev/tools/cerebras) | 56.5 | 65% | Wafer-scale inference — the fastest tokens per second available |
| 18 | [DeepSeek](https://toolweight.dev/tools/deepseek) | 46.2 | 83% | Frontier-adjacent models at a small fraction of Western prices |

Re-rank against your own priorities by appending `?w=field_id:weight,…` to the canonical URL, or call `score_tools` on the MCP server at `https://toolweight.dev/mcp`.

## Comparison

### Capability

| Tool | Flagship model | Context window (tokens) | Max output (tokens) | Image input | Audio in/out | Open weights | OpenAI-compat API |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Google Gemini | Gemini 3 Pro | 1,000,000 tokens | — | ● | ● | ◐ | ● |
| OpenAI | GPT-5.1 | 272,000 tokens | 128,000 tokens | ● | ● | ◐ | ● |
| Amazon Bedrock | Multi-vendor — Claude Opus 4.8, Llama 4, Mistral, Nova Premier | 1,000,000 tokens | 128,000 tokens | ● | ◐ | ◐ | ○ |
| Mistral AI | Mistral Large 3 | — | — | ● | ◐ | ◐ | ● |
| Qwen | Qwen3-Max | 262,144 tokens | — | ◐ | ◐ | ◐ | ● |
| Meta Llama | Llama 4 Maverick | 1,000,000 tokens | — | ● | ○ | ● | ● |
| Self-hosted (vLLM) | Whatever you deploy — DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4 | — | — | ◐ | ○ | ● | ● |
| Together AI | Open-weight catalogue — DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4 | — | — | ◐ | ◐ | ● | ● |
| xAI | Grok 4.1 | — | — | ● | ○ | ◐ | ● |
| Fireworks AI | Open-weight catalogue — DeepSeek, Kimi K2, Qwen3, Llama 4 | — | — | ◐ | ◐ | ● | ● |
| Cohere | Command A (command-a-03-2025) | 256,000 tokens | — | ◐ | ○ | ◐ | ◐ |
| Anthropic | Claude Fable 5 (claude-fable-5) | 1,000,000 tokens | 128,000 tokens | ● | ○ | ○ | ◐ |
| OpenRouter | Whichever upstream model you route to (400+ available) | — | — | ◐ | ◐ | ◐ | ● |
| Z.ai (GLM) | GLM-4.6 | 200,000 tokens | — | ◐ | ○ | ● | ● |
| Moonshot AI | Kimi K2 Thinking | 256,000 tokens | — | ○ | ○ | ● | ● |
| Groq | Curated open-weight models (Kimi K2, Llama 4, Qwen3, GPT-OSS) | — | — | ◐ | ◐ | ● | ● |
| Cerebras | Curated open-weight models (Qwen3, GLM, Llama, GPT-OSS) | — | — | ○ | ○ | ● | ● |
| DeepSeek | DeepSeek-V3.2 (deepseek-chat / deepseek-reasoner) | 128,000 tokens | — | ○ | ○ | ● | ● |

### Performance

| Tool | TTFT p50 | Output tok/s (tok/s) |
| --- | --- | --- |
| Google Gemini | — | — |
| OpenAI | — | — |
| Amazon Bedrock | — | — |
| Mistral AI | — | — |
| Qwen | — | — |
| Meta Llama | — | — |
| Self-hosted (vLLM) | — | — |
| Together AI | — | — |
| xAI | — | — |
| Fireworks AI | — | — |
| Cohere | — | — |
| Anthropic | — | — |
| OpenRouter | — | — |
| Z.ai (GLM) | — | — |
| Moonshot AI | — | — |
| Groq | 250 ms | 400 tok/s |
| Cerebras | 200 ms | 2,000 tok/s |
| DeepSeek | — | — |

### Pricing

| Tool | $/M input (/M tok) | $/M output (/M tok) | $/M cache read (/M tok) | Batch discount (%) |
| --- | --- | --- | --- | --- |
| Google Gemini | $2 /M tok | $12 /M tok | $0.2 /M tok | 50 % |
| OpenAI | $1.25 /M tok | $10 /M tok | $0.125 /M tok | 50 % |
| Amazon Bedrock | — | — | — | 50 % |
| Mistral AI | — | — | — | 50 % |
| Qwen | $1.2 /M tok | $6 /M tok | — | 50 % |
| Meta Llama | — | — | — | — |
| Self-hosted (vLLM) | — | — | — | 0 % |
| Together AI | — | — | — | 50 % |
| xAI | — | — | — | — |
| Fireworks AI | — | — | — | — |
| Cohere | $2.5 /M tok | $10 /M tok | — | — |
| Anthropic | $10 /M tok | $50 /M tok | $1 /M tok | 50 % |
| OpenRouter | — | — | — | 0 % |
| Z.ai (GLM) | $0.6 /M tok | $2.2 /M tok | — | — |
| Moonshot AI | $0.6 /M tok | $2.5 /M tok | $0.15 /M tok | — |
| Groq | — | — | — | — |
| Cerebras | — | — | — | — |
| DeepSeek | $0.28 /M tok | $0.42 /M tok | $0.028 /M tok | 0 % |

### Agentic

| Tool | Tool use | Schema output | Effort control | Computer use | MCP support | Cache TTL |
| --- | --- | --- | --- | --- | --- | --- |
| Google Gemini | ● | ● | ● | ◐ | ◐ | Explicit caches, default 1 h TTL; implicit caching too |
| OpenAI | ● | ● | ● | ◐ | ● | Automatic, ~5–60 min, no configuration |
| Amazon Bedrock | ● | ● | ● | ● | ○ | 5 min default, 1 h option; explicit breakpoints only |
| Mistral AI | ● | ● | ◐ | ○ | ◐ | — |
| Qwen | ● | ◐ | ◐ | ○ | ◐ | — |
| Meta Llama | ◐ | ◐ | ○ | ○ | ○ | Host-dependent |
| Self-hosted (vLLM) | ◐ | ● | ○ | ○ | ○ | Automatic prefix cache in VRAM, no TTL |
| Together AI | ◐ | ● | ○ | ○ | ○ | — |
| xAI | ● | ● | ◐ | ○ | ○ | — |
| Fireworks AI | ◐ | ● | ○ | ○ | ○ | — |
| Cohere | ● | ● | ○ | ○ | ○ | — |
| Anthropic | ● | ● | ● | ● | ● | 5 min default, 1 h option; automatic or explicit breakpoints |
| OpenRouter | ● | ◐ | ◐ | ○ | ○ | Passed through where the upstream supports caching |
| Z.ai (GLM) | ● | ◐ | ◐ | ○ | ○ | — |
| Moonshot AI | ● | ◐ | ◐ | ○ | ○ | Explicit context caching, paid per storage hour |
| Groq | ◐ | ◐ | ○ | ○ | ○ | — |
| Cerebras | ◐ | ◐ | ○ | ○ | ○ | — |
| DeepSeek | ◐ | ◐ | ◐ | ○ | ○ | Automatic disk cache, no configuration |

### Governance & continuity

| Tool | Continuity policy | Notice period (days) | Pinnable versions | Zero retention | Trains on your data | AU region |
| --- | --- | --- | --- | --- | --- | --- |
| Google Gemini | Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not. | — | ● | ◐ | ◐ | ● |
| OpenAI | Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's — closer to a quarter than half a year for API snapshots — and preview models have been pulled faster still. | 90 days | ● | ◐ | ○ | — |
| Amazon Bedrock | The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do. | — | ● | ● | ○ | ● |
| Mistral AI | Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away. | — | ● | ◐ | ○ | ○ |
| Qwen | Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely. | — | ● | — | ◐ | ◐ |
| Meta Llama | Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you — if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears. | — | ● | ◐ | ○ | ● |
| Self-hosted (vLLM) | Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster. | — | ● | ● | ○ | ● |
| Together AI | Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint. | — | ● | — | ○ | ○ |
| xAI | Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here. | — | ● | — | ◐ | ○ |
| Fireworks AI | Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient. | — | ● | — | ○ | ○ |
| Cohere | Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does. | — | ● | ◐ | ○ | ◐ |
| Anthropic | Published deprecation page lists a retirement date per model, typically months ahead — Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee. | 180 days | ◐ | ◐ | ○ | ◐ |
| OpenRouter | Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself — a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement. | — | ◐ | ◐ | ○ | ○ |
| Z.ai (GLM) | Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model. | — | ● | ○ | ◐ | ○ |
| Moonshot AI | Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence — download the checkpoint and continuity is your problem, not theirs. | — | ● | ○ | ◐ | ○ |
| Groq | The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun. | — | ◐ | — | ○ | ○ |
| Cerebras | Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite. | — | ◐ | — | ○ | ○ |
| DeepSeek | The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases — deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently. | — | ○ | ○ | ● | ○ |

### Traction

| Tool | API since |
| --- | --- |
| Google Gemini | 2023 |
| OpenAI | 2020 |
| Amazon Bedrock | 2023 |
| Mistral AI | 2023 |
| Qwen | 2023 |
| Meta Llama | 2023 |
| Self-hosted (vLLM) | 2023 |
| Together AI | 2023 |
| xAI | 2024 |
| Fireworks AI | 2023 |
| Cohere | 2021 |
| Anthropic | 2023 |
| OpenRouter | 2023 |
| Z.ai (GLM) | 2023 |
| Moonshot AI | 2023 |
| Groq | 2024 |
| Cerebras | 2024 |
| DeepSeek | 2023 |

### Positioning

| Tool | Positioning |
| --- | --- |
| Google Gemini | Long context and native multimodality across Google's cloud |
| OpenAI | One API for text, reasoning, audio, images and embeddings |
| Amazon Bedrock | Frontier models inside your existing AWS security perimeter |
| Mistral AI | European models with open weights and deployable anywhere |
| Qwen | The widest open-weight family, plus a closed flagship tier |
| Meta Llama | Open-weight models you can host anywhere, forever |
| Self-hosted (vLLM) | Your weights, your GPUs, your OpenAI-compatible endpoint |
| Together AI | Open-weight models with a Western contract and dedicated capacity |
| xAI | Fast, cheap frontier models with live access to X |
| Fireworks AI | Low-latency open-weight inference with strict structured output |
| Cohere | Enterprise RAG models you can deploy inside your own network |
| Anthropic | Frontier models for long-horizon agentic work and code |
| OpenRouter | One key, one API shape, every model worth calling |
| Z.ai (GLM) | Coding-focused GLM models with a flat-rate subscription |
| Moonshot AI | Open-weight agentic performance at open-weight prices |
| Groq | Custom LPU silicon for open-weight models at very low latency |
| Cerebras | Wafer-scale inference — the highest tokens per second available |
| DeepSeek | Near-frontier quality at a fraction of the token price |

## Field registry

| Field | id | Group | Type | Unit | Better | Default weight | Definition |
| --- | --- | --- | --- | --- | --- | --- | --- |
| Flagship model | `flagship_model` | capability | text | — | n/a | 0 | The provider's current top-end model, and the model every price, context and capability figure on this row describes. |
| Context window | `context_window` | capability | number | tokens | higher | 5 | Maximum INPUT tokens the flagship model accepts in a single request, at the standard (non-surcharged) tier. Output is capped separately — see Max output. Where a vendor advertises one combined figure, this column carries the input half. |
| Max output | `max_output` | capability | number | tokens | higher | 3 | Maximum tokens the flagship model will emit in a single response, counted separately from the input context window, and inclusive of reasoning tokens where those are billed as output. Blank where the provider publishes no figure for the named flagship, or where the row is a catalogue rather than a model. |
| Image input | `vision` | capability | tristate | — | higher | 4 | Whether the flagship model accepts images in the same request as text; 'partial' means images work but at reduced resolution or only on a sibling model. |
| Audio in/out | `audio_io` | capability | tristate | — | higher | 3 | Native speech input and speech output through the same API surface; 'partial' means one direction only, or a separate speech model rather than native audio tokens. |
| Open weights | `open_weights` | capability | tristate | — | higher | 4 | Whether you can download and self-host weights of comparable capability from this vendor; 'partial' means smaller or older models are released but the flagship is closed. |
| OpenAI-compat API | `openai_compatible` | capability | tristate | — | higher | 5 | Whether the provider serves a /v1/chat/completions endpoint the OpenAI SDKs can talk to unmodified; 'partial' means a compatibility shim exists but omits tools, streaming or structured output. |
| TTFT p50 | `ttft_p50` | performance | duration | — | lower | 4 | Median time from request to first streamed token on a short prompt, as reported by third-party benchmarks; blank where no credible published figure exists. |
| Output tok/s | `throughput_tps` | performance | number | tok/s | higher | 4 | Sustained single-stream output tokens per second on the named flagship model, community-measured; blank where unpublished. |
| $/M input | `price_in` | pricing | currency | /M tok | lower | 8 | Pay-as-you-go list price per million input tokens for the named flagship model, uncached, at the standard context tier. |
| $/M output | `price_out` | pricing | currency | /M tok | lower | 9 | Pay-as-you-go list price per million output tokens for the named flagship model, including reasoning or thinking tokens where those are billed as output. |
| $/M cache read | `cached_input` | pricing | currency | /M tok | lower | 6 | Price per million input tokens served from a prompt cache hit; excludes any separate cache-write premium or per-hour cache storage fee. |
| Batch discount | `batch_discount` | pricing | number | % | higher | 4 | Percentage off list price for asynchronous batch submission with a multi-hour completion window; 0 means no batch endpoint is offered. |
| Tool use | `tool_use` | agentic | tristate | — | higher | 8 | Native function calling with parallel calls in a single assistant turn; 'partial' means tools work but serially, or degrade noticeably over a long agent loop. |
| Schema output | `structured_output` | agentic | tristate | — | higher | 7 | Constrained decoding that guarantees the response validates against a supplied JSON Schema; 'partial' means JSON mode without schema enforcement. |
| Effort control | `reasoning_effort` | agentic | tristate | — | higher | 6 | A request-level parameter that trades reasoning depth against cost and latency; 'partial' means thinking can only be switched on or off, with no graduated levels. |
| Computer use | `computer_use` | agentic | tristate | — | higher | 3 | A screenshot-and-act tool for driving a GUI. 'yes' means the row's named flagship model drives it directly and returns coordinates itself — a beta header is fine, since the whole category ships this behind one. 'partial' means it is delivered through a separate specialised or preview model rather than the flagship, or is gated behind a restricted preview. |
| MCP support | `mcp_support` | agentic | tristate | — | higher | 5 | Whether the API itself can connect to Model Context Protocol servers, rather than leaving MCP entirely to your client; 'partial' means SDK-side support only. |
| Cache TTL | `prompt_cache_ttl` | agentic | text | — | n/a | 0 | How long a cached prompt prefix stays warm, and whether caching is automatic or requires explicit breakpoints in the request. |
| Continuity policy | `continuity_policy` | governance | markdown | — | n/a | 0 | The killer field: how much notice you get before a model is retired, and how long a pinned version keeps serving after the alias moves on. |
| Notice period | `deprecation_notice` | governance | number | days | higher | 7 | Typical days between a model being marked deprecated and the endpoint returning 404, from published policy or observed retirements; blank where the provider publishes nothing. |
| Pinnable versions | `pinned_snapshots` | governance | tristate | — | higher | 7 | Whether you can address an immutable dated model version that will not change behaviour underneath you; 'partial' means some models are pinnable and the flagship is not. |
| Zero retention | `zero_retention` | governance | tristate | — | higher | 6 | Whether prompts and completions can be excluded from all storage, including abuse-monitoring logs; 'partial' means available on request, on a higher tier, or with model-level exceptions. |
| Trains on your data | `trains_on_data` | governance | tristate | — | lower | 7 | Whether API inputs and outputs are used to train models by default; 'partial' means opt-out is available but the default is on, or the terms are ambiguous. |
| AU region | `au_region` | governance | tristate | — | higher | 4 | Whether inference can be pinned to Australian infrastructure; 'partial' means data residency controls exist for some regions but AU is unconfirmed or indirect. |
| API since | `launch_year` | traction | number | — | lower | 1 | Year the provider first made a general-availability inference API public, as a rough proxy for operational maturity. |
| Positioning | `positioning` | marketing | text | — | n/a | 0 | How the provider positions itself in one line, paraphrased from its own homepage rather than quoted. |

## Presets

| Preset | id | What it optimises for | Weights | Ranked URL |
| --- | --- | --- | --- | --- |
| Cheapest per token | `cheapest-per-token` | Ranks on flagship list price, with cache and batch discounts weighted in. Rows that publish no flat per-token rate for a named flagship — the routers, the multi-model catalogues and the self-host baseline — are marked inapplicable rather than ranked, because 'cheapest' is not a question they answer. | price_in:10, price_out:10, cached_input:6, batch_discount:4, context_window:2, open_weights:2 | https://toolweight.dev/compare/llm-apis?w=price_in:10,price_out:10,cached_input:6,batch_discount:4,context_window:2,open_weights:2 |
| Fastest TTFT | `fastest-ttft` | Time to first token and sustained throughput. Deliberately sparse: only the providers whose entire pitch is speed have a credible published figure, so everyone else is marked inapplicable rather than flattered by having no number at all. | ttft_p50:10, throughput_tps:8, price_out:2, context_window:1 | https://toolweight.dev/compare/llm-apis?w=ttft_p50:10,throughput_tps:8,price_out:2,context_window:1 |
| Best tool use | `best-tool-use` | For agent harnesses: tool calling, schema-guaranteed output, effort control and MCP. | tool_use:10, structured_output:9, mcp_support:7, reasoning_effort:6, computer_use:4, context_window:3 | https://toolweight.dev/compare/llm-apis?w=tool_use:10,structured_output:9,mcp_support:7,reasoning_effort:6,computer_use:4,context_window:3 |
| Most privacy-preserving | `most-private` | Zero data retention, no training on your prompts, and the option to run the weights yourself. | zero_retention:10, trains_on_data:10, open_weights:6, au_region:5, pinned_snapshots:2 | https://toolweight.dev/compare/llm-apis?w=zero_retention:10,trains_on_data:10,open_weights:6,au_region:5,pinned_snapshots:2 |
| Longest continuity | `longest-continuity` | Will the app still work in eighteen months? Notice period, pinnable snapshots, escape hatches. | deprecation_notice:10, pinned_snapshots:10, open_weights:8, openai_compatible:6 | https://toolweight.dev/compare/llm-apis?w=deprecation_notice:10,pinned_snapshots:10,open_weights:8,openai_compatible:6 |

## Verdict

Every provider on this page will quote you a $/M token figure. Almost none of them will tell you how long the model behind it lives. That asymmetry is the single biggest hidden cost in this category, and it is why this page leads with a continuity column rather than a price column.

For most production work in 2026 the shortlist is short. Anthropic if the workload is agentic — tool calls, long multi-step runs, code — because effort control, prompt caching and MCP are first-class rather than bolted on. OpenAI if you need audio, image generation and text behind one billing relationship, or if your team is already fluent in the Responses API. Google if you need million-token context cheaply, native audio in and out, or an AU-resident deployment through Vertex. Those three are not interchangeable at the prompt level: a prompt tuned on one lands differently on the others, and the effort/thinking knobs have no common vocabulary.

The Chinese open-weight labs — DeepSeek, Moonshot, Qwen, Z.ai — have changed what the floor looks like. DeepSeek's flagship bills $0.42 per million output tokens against $50 for Claude Fable 5 and $15 for Claude Sonnet 5 — roughly a hundred and twenty times cheaper than Anthropic's top-end model, and still around thirty-five times cheaper than the volume tier most teams would actually reach for. For classification, extraction, summarisation and bulk rewriting the quality gap does not justify the price gap. The catch is governance, not capability: retention defaults are permissive, zero-data-retention is generally not on offer, and several expose only a rolling alias that silently upgrades under you. If the data is sensitive, run the weights yourself or route through a host that is contractually clean rather than calling the origin API.

Routers deserve more credit than they get. OpenRouter is the cheapest insurance policy in the category — one integration, hundreds of models, and a real chance that a retired model stays reachable somewhere because an open-weight copy is still hosted. Groq and Cerebras are not general-purpose substitutes; they are latency instruments. If your product's differentiator is that the answer appears instantly — voice, autocomplete, interactive search — they are worth an entire architecture. If it is not, their model menu will frustrate you within a quarter.

Where this is heading: prices keep falling, context windows have stopped being the differentiator, and the fight has moved to agentic fidelity — parallel tool calls that do not degrade, structured output that never breaks schema, caching that survives a long agent loop. Assume the model you ship on today will be retired inside two years. Build the abstraction layer now, keep an evaluation set you can re-run in an afternoon, and treat every provider on this page as replaceable.

## The pricing traps nobody puts on the pricing page

Headline $/M is the least useful number in this category. Four things distort it.

First, output tokens dominate. A reasoning model can burn tens of thousands of thinking tokens before it writes a word you see, and those bill at the output rate. A provider with a cheap input rate and an expensive output rate is expensive for reasoning work and cheap for retrieval-augmented answering. Model your actual input:output ratio before comparing.

Second, cached input is where the real money is. Cache reads bill at roughly a tenth of the input rate at Anthropic and OpenAI, and an agent loop resending its history every turn is mostly cache reads. But caching is a prefix match — a timestamp in the system prompt, a non-deterministic JSON serialisation, or a tool list that varies per user destroys it silently. Watch the cache-read token counter, not the invoice.

Third, long-context surcharges. Google charges a higher rate above 200K tokens on its flagship; others quietly meter cache storage per hour. A million-token window advertised at the short-context rate is not what you will pay.

Fourth, batch — but check the column before you budget for it. Where an asynchronous endpoint exists the discount is usually a flat 50%, and Anthropic, OpenAI, Google, Mistral and Qwen all publish that rate, as do Bedrock and Together. It is not universal. DeepSeek offers no batch endpoint at all — its list price already sits below most competitors' batch rates — and neither does OpenRouter, which is a real cost if half your workload could tolerate async. xAI, Moonshot, Z.ai, Cohere, Fireworks and Cerebras publish nothing we could confirm, and Groq's batch discount is documented as existing without a percentage we could stand behind. Where the halving is there and part of your workload tolerates a few hours of latency, it is the one saving no negotiation will beat; where it is not, no amount of committed spend conjures it.

## How to make a migration survivable

Assume an eighteen-month life for whatever you ship on. The providers that publish retirement dates are doing you a favour; the ones that don't are the ones that will break you.

Pin a dated snapshot wherever the provider offers one, and record the pin in configuration rather than code. Note that this is getting harder, not easier: Anthropic's current line-up ships as bare aliases with no dated model ID behind them, so 'pin the snapshot' is not always an available move any more, and you fall back on the published retirement date instead.

Keep the provider behind one seam. Not an abstraction that pretends every model is the same — that fails the moment you touch thinking blocks or tool-result shapes — but a single module that owns request construction, retry policy and response parsing per provider. Two implementations behind one interface is the cheap insurance; a universal adapter is the expensive fantasy.

Maintain an evaluation set you can re-run in an afternoon: fifty to two hundred real inputs with graded expected outputs. Without it, a forced migration becomes a multi-week vibe check. With it, a migration is a morning's work and a number you can defend.

Finally, know your escape hatch before you need it. For open-weight models it is total — the weights are on disk and vLLM will serve them for as long as you have GPUs. For closed models it is a router that may still have an upstream, or nothing at all.

## Deciding: three workloads, three answers

Bulk text processing — classification, extraction, tagging, translation, summarising a firehose. Price dominates and capability differences barely register. DeepSeek, Qwen, Moonshot and Z.ai are the rational choices, or Groq and Together if you want open weights with a Western contract. Add the batch endpoint if latency is negotiable. Do not put sensitive data through a provider whose retention terms you haven't read.

Agent harnesses — a loop that plans, calls tools, reads results and iterates for minutes at a time. This is where the frontier labs earn their price. You need tool calls that stay well-formed at depth, structured output that is enforced rather than requested, an effort or thinking control so you can trade cost against quality per route, and prompt caching that survives the loop. Anthropic and OpenAI are the credible options; Google is close and cheaper on long context. Budget for the fact that a long agentic turn can now run for minutes on a single request.

Interactive, latency-critical UI — voice, live search, inline completion. Time to first token is the product. Cerebras and Groq are in a different class here, an order of magnitude ahead of first-party endpoints on tokens per second, and the constraint is that you take whichever open-weight models they have chosen to host. Design the feature around the model menu, not the other way around.

## Provenance — every cell, every source

### Google Gemini {#google-gemini}

Gemini via AI Studio for prototyping or Vertex AI for production

|  |  |
| --- | --- |
| Site | https://ai.google.dev |
| Docs | https://ai.google.dev/gemini-api/docs |
| Company | Google |
| Founded | 1998 |
| Open source | No |
| Profile | https://toolweight.dev/tools/google-gemini |
| Score (default weights) | 79.9 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Gemini 3 Pro | Inferred | 2026-01-15 | https://blog.google/products/gemini/gemini-3/ | Launched November 2025; a newer flagship may have shipped since. | #google-gemini-flagship_model |
| Context window | 1,000,000 tokens | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/models | Input window. Google publishes a separate, much smaller output cap alongside it. | #google-gemini-context_window |
| Max output | — | Unknown | — | — | Google publishes a per-model output cap well below the million-token input window, but the current figure for this flagship was not confirmed and is not worth guessing at — check the models page before sizing a long generation. | #google-gemini-max_output |
| Image input | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/image-understanding | — | #google-gemini-vision |
| Audio in/out | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/live | Native audio tokens in and out via the Live API, plus video input. | #google-gemini-audio_io |
| Open weights | ◐ | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemma | The Gemma family is downloadable; Gemini itself is not. | #google-gemini-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/openai | A chat-completions compatibility endpoint is documented for both AI Studio and Vertex. | #google-gemini-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #google-gemini-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #google-gemini-throughput_tps |
| $/M input | $2 /M tok | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/pricing | Rate for prompts up to 200K tokens, off Google's pricing page; longer prompts are surcharged. | #google-gemini-price_in |
| $/M output | $12 /M tok | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/pricing | Rate for prompts up to 200K tokens; the long-context tier costs more. | #google-gemini-price_out |
| $/M cache read | $0.2 /M tok | Inferred | 2026-01-15 | — | Approximate — caching discounts input heavily but Google also charges cache storage per token-hour, which can dominate for small caches held a long time. | #google-gemini-cached_input |
| Batch discount | 50 % | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/batch-mode | — | #google-gemini-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/function-calling | — | #google-gemini-tool_use |
| Schema output | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/structured-output | responseSchema constrains decoding to a supplied schema. | #google-gemini-structured_output |
| Effort control | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/thinking | Thinking level / thinking budget controls reasoning depth per request. | #google-gemini-reasoning_effort |
| Computer use | ◐ | Inferred | 2026-01-15 | https://ai.google.dev/gemini-api/docs/computer-use | A dedicated computer-use model rather than a tool on the flagship — which is this column's definition of partial, so it is now graded that way rather than as a yes. | #google-gemini-computer_use |
| MCP support | ◐ | Inferred | 2026-01-15 | — | MCP is wired up in the SDKs rather than being a server-side connector on the API. | #google-gemini-mcp_support |
| Cache TTL | Explicit caches, default 1 h TTL; implicit caching too | Inferred | 2026-01-15 | — | — | #google-gemini-prompt_cache_ttl |
| Continuity policy | Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not. | Community-reported | 2026-01-15 | — | — | #google-gemini-continuity_policy |
| Notice period | — | Unknown | — | — | Google publishes per-version retirement dates but no single guaranteed notice window; preview models get materially less than stable ones. | #google-gemini-deprecation_notice |
| Pinnable versions | ● | Vendor-claimed | 2026-01-15 | https://ai.google.dev/gemini-api/docs/models | — | #google-gemini-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-01-15 | — | Tier-dependent, which is what this column calls partial: on Vertex, prompts are not stored and enterprise controls apply, but the free AI Studio tier is a different bargain — assume it is used to improve products. We have not confirmed exclusion from abuse-monitoring logs, which is the bar the column's definition actually sets. | #google-gemini-zero_retention |
| Trains on your data | ◐ | Inferred | 2026-01-15 | — | No for paid Vertex and paid AI Studio usage; the free tier explicitly is used to improve Google products. | #google-gemini-trains_on_data |
| AU region | ● | Inferred | 2026-01-15 | — | Vertex AI serves Gemini from australia-southeast1, with region pinning enforced at the project level. | #google-gemini-au_region |
| API since | 2023 | Inferred | — | — | — | #google-gemini-launch_year |
| Positioning | Long context and native multimodality across Google's cloud | Inferred | — | — | — | #google-gemini-positioning |

**Verdict.** The cheapest credible route to a million-token window and the only provider here with genuinely native audio in and out. Vertex is also the most convincing answer to an Australian data-residency requirement. The cost is churn — this line-up turns over faster than any other, so pin versions and expect a migration each year.

**Pick it when**

- Workloads that genuinely need 500K+ tokens of context per request
- Voice and video products wanting one model rather than a pipeline
- Australian or regulated deployments that must pin inference to a region

### OpenAI {#openai}

GPT models plus audio, images and embeddings on one bill

|  |  |
| --- | --- |
| Site | https://openai.com |
| Docs | https://platform.openai.com/docs |
| Company | OpenAI |
| Founded | 2015 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/openai |
| Score (default weights) | 79.8 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | GPT-5.1 | Inferred | 2026-01-15 | — | Last confirmed January 2026; OpenAI ships flagships faster than this page's cadence, so verify the current top model before quoting. | #openai-flagship_model |
| Context window | 272,000 tokens | Inferred | 2026-01-15 | — | OpenAI advertises 400K as a combined figure; that is 272K of input plus a 128K output cap. This column carries the input half so it compares like for like against the input-only windows every other row publishes. | #openai-context_window |
| Max output | 128,000 tokens | Inferred | 2026-01-15 | — | The output half of the advertised 400K combined window. | #openai-max_output |
| Image input | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/images-vision | — | #openai-vision |
| Audio in/out | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/realtime | Realtime API does speech-to-speech; separate transcription and TTS models are also available. | #openai-audio_io |
| Open weights | ◐ | Vendor-claimed | 2026-01-15 | https://openai.com/index/introducing-gpt-oss/ | gpt-oss-120b and gpt-oss-20b released under Apache 2.0 in August 2025. The flagship remains closed. | #openai-open_weights |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | — | It is the reference implementation, though the newer Responses API is where the agentic features live. | #openai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #openai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #openai-throughput_tps |
| $/M input | $1.25 /M tok | Vendor-claimed | 2026-01-15 | https://openai.com/api/pricing/ | GPT-5 family list rate as of January 2026, read off OpenAI's own pricing page. The flagship name above it is the inferred part of this row, not the number. | #openai-price_in |
| $/M output | $10 /M tok | Vendor-claimed | 2026-01-15 | https://openai.com/api/pricing/ | — | #openai-price_out |
| $/M cache read | $0.125 /M tok | Inferred | 2026-01-15 | — | Automatic caching discounts repeated prefixes by roughly 90%; no request-side configuration. | #openai-cached_input |
| Batch discount | 50 % | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/batch | — | #openai-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/function-calling | — | #openai-tool_use |
| Schema output | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/structured-outputs | Strict JSON Schema enforcement on both responses and tool parameters. | #openai-structured_output |
| Effort control | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/reasoning | reasoning_effort parameter with graduated levels. | #openai-reasoning_effort |
| Computer use | ◐ | Inferred | 2026-01-15 | https://platform.openai.com/docs/guides/tools-computer-use | The capability exists and is documented, but it is delivered through a separate computer-use preview model rather than by the flagship named on this row — the column's definition of partial. Graded on the same basis as Google's dedicated computer-use model. | #openai-computer_use |
| MCP support | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/tools-remote-mcp | Remote MCP servers can be attached as a tool on the Responses API. | #openai-mcp_support |
| Cache TTL | Automatic, ~5–60 min, no configuration | Inferred | 2026-01-15 | — | — | #openai-prompt_cache_ttl |
| Continuity policy | Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's — closer to a quarter than half a year for API snapshots — and preview models have been pulled faster still. | Inferred | 2026-01-15 | — | — | #openai-continuity_policy |
| Notice period | 90 days | Inferred | — | — | Observed from past API snapshot retirements; not a published guarantee, and varies by model tier. | #openai-deprecation_notice |
| Pinnable versions | ● | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/models | — | #openai-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-01-15 | — | Default abuse-monitoring retention is 30 days; zero-retention is available to approved customers. | #openai-zero_retention |
| Trains on your data | ○ | Vendor-claimed | 2026-01-15 | https://platform.openai.com/docs/guides/your-data | API data has not been used for training by default since March 2023. | #openai-trains_on_data |
| AU region | — | Unknown | — | — | Data residency is offered in several regions; AU availability unconfirmed at the time of writing. | #openai-au_region |
| API since | 2020 | Inferred | — | — | — | #openai-launch_year |
| Positioning | One API for text, reasoning, audio, images and embeddings | Inferred | — | — | — | #openai-positioning |

**Verdict.** The widest surface area in the category, and the only one where audio, images and text share a bill and a mental model. Dated snapshots make routine version management painless — the strongest routine protection here. On notice periods, be careful how much weight you put on the comparison: the retirements we have observed have run shorter than Anthropic's, but neither provider publishes a guaranteed minimum, so both numbers in that column are inference from past behaviour rather than a commitment either vendor has made. Pin, and keep an eye on the deprecations page.

**Pick it when**

- Multimodal products that need speech and image generation alongside text
- Teams standardising on one vendor for everything rather than best-of-breed
- Anyone who wants a dated snapshot they can actually pin

### Amazon Bedrock {#amazon-bedrock}

Multi-vendor model access inside your existing AWS account

|  |  |
| --- | --- |
| Site | https://aws.amazon.com/bedrock/ |
| Docs | https://docs.aws.amazon.com/bedrock/ |
| Company | AWS |
| Founded | 2006 |
| Open source | No |
| Profile | https://toolweight.dev/tools/amazon-bedrock |
| Score (default weights) | 78.1 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Multi-vendor — Claude Opus 4.8, Llama 4, Mistral, Nova Premier | Inferred | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html | Bedrock is a catalogue, not a model, so there is no vendor claim to make about a flagship. The models listed are a toolweight selection from the supported-models page — an editorial choice about what the catalogue is for, which is why this cell is graded inferred rather than vendor-claimed. Capability figures below describe Claude on Bedrock as the reference deployment except where a cell says otherwise. | #amazon-bedrock-flagship_model |
| Context window | 1,000,000 tokens | Inferred | 2026-06-24 | — | Input window for Claude on Bedrock; varies by model. | #amazon-bedrock-context_window |
| Max output | 128,000 tokens | Inferred | 2026-06-24 | — | For Claude on Bedrock, tracking the first-party output cap; varies by model across the catalogue. | #amazon-bedrock-max_output |
| Image input | ● | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/conversation-inference-supported-models-features.html | — | #amazon-bedrock-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | Amazon's own Nova Sonic does speech-to-speech, but this row's reference deployment is Claude on Bedrock, which has no audio at all. Regraded to partial because the capability exists somewhere in the catalogue rather than on the model every other figure in this row describes — a catalogue capability, not a flagship one. | #amazon-bedrock-audio_io |
| Open weights | ◐ | Inferred | 2026-01-15 | — | Hosts open-weight models (Llama, Mistral) alongside closed ones; publishes none of its own. | #amazon-bedrock-open_weights |
| OpenAI-compat API | ○ | Inferred | 2026-06-24 | — | Uses the AWS SDK and SigV4 auth. Anthropic ships a dedicated Bedrock client rather than a base-URL swap. | #amazon-bedrock-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #amazon-bedrock-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #amazon-bedrock-throughput_tps |
| $/M input | — | Unknown | — | — | Set by AWS per model and not guaranteed to match first-party rates; Claude on Bedrock has historically tracked close to Anthropic list price. | #amazon-bedrock-price_in |
| $/M output | — | Unknown | — | — | — | #amazon-bedrock-price_out |
| $/M cache read | — | Unknown | — | — | — | #amazon-bedrock-cached_input |
| Batch discount | 50 % | Inferred | 2026-01-15 | — | Bedrock batch inference is offered at a reduced rate for supported models. | #amazon-bedrock-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/tool-use.html | — | #amazon-bedrock-tool_use |
| Schema output | ● | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/structured-output.html | Available for Claude on Bedrock; support varies by model. | #amazon-bedrock-structured_output |
| Effort control | ● | Inferred | 2026-06-24 | — | — | #amazon-bedrock-reasoning_effort |
| Computer use | ● | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/computer-use.html | Beta for Claude on Bedrock — the same Claude computer-use beta the first-party row carries, driven by the same flagship model. Upgraded from partial so the two rows are graded on one rule: beta status alone does not make a yes into a partial, or Anthropic's own row would have to move too. | #amazon-bedrock-computer_use |
| MCP support | ○ | Inferred | 2026-06-24 | — | The MCP connector is not available on Bedrock; wire MCP up client-side instead. | #amazon-bedrock-mcp_support |
| Cache TTL | 5 min default, 1 h option; explicit breakpoints only | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html | Both TTLs are available, same as first-party — it is only the automatic top-level caching path that Bedrock does not offer, so every cache boundary here has to be placed by hand. | #amazon-bedrock-prompt_cache_ttl |
| Continuity policy | The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do. | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html | — | #amazon-bedrock-continuity_policy |
| Notice period | — | Unknown | — | — | AWS publishes per-model EOL dates through the lifecycle process; no single fixed minimum window was confirmed. | #amazon-bedrock-deprecation_notice |
| Pinnable versions | ● | Vendor-claimed | 2026-06-24 | https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html | Versioned model IDs and inference-profile ARNs. | #amazon-bedrock-pinned_snapshots |
| Zero retention | ● | Vendor-claimed | 2026-01-15 | https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html | AWS does not store prompts or completions; model invocation logging is opt-in and lands in your own account. This stays a yes where Google's and Cohere's moved to partial because it is a documented, unconditional vendor statement rather than a property of one tier or one deployment mode. | #amazon-bedrock-zero_retention |
| Trains on your data | ○ | Vendor-claimed | 2026-01-15 | https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html | — | #amazon-bedrock-trains_on_data |
| AU region | ● | Inferred | 2026-01-15 | — | ap-southeast-2 (Sydney) serves frontier models; check per-model availability as it is not uniform across regions. | #amazon-bedrock-au_region |
| API since | 2023 | Inferred | — | — | — | #amazon-bedrock-launch_year |
| Positioning | Frontier models inside your existing AWS security perimeter | Inferred | — | — | — | #amazon-bedrock-positioning |

**Verdict.** Choose Bedrock for procurement and governance, not capability. IAM, VPC endpoints, an existing contract, a Sydney region and a published model lifecycle are worth real money to regulated teams. Accept that you will be weeks or months behind on features, and that the MCP connector and automatic caching are simply not there.

**Pick it when**

- Regulated workloads that must stay inside an existing AWS perimeter
- Australian data residency without self-hosting
- Teams that need a documented model lifecycle to satisfy an auditor

### Mistral AI {#mistral-ai}

European lab with an open-weight lineage and EU-resident hosting

|  |  |
| --- | --- |
| Site | https://mistral.ai |
| Docs | https://docs.mistral.ai |
| Company | Mistral AI |
| Founded | 2023 |
| Funding | Late-stage private |
| Open source | Yes (Apache-2.0 (selected models)) |
| Profile | https://toolweight.dev/tools/mistral-ai |
| Score (default weights) | 72.8 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Mistral Large 3 | Inferred | 2026-01-15 | — | Mistral Medium is the better price-performance pick for most workloads on this platform. | #mistral-ai-flagship_model |
| Context window | — | Unknown | — | — | Withdrawn. The 128K figure previously here is the Mistral Large 2 (mistral-large-2411) window, not Large 3's — Large 3 shipped with a materially larger one. Blank until the flagship's own figure is verified against Mistral's docs. | #mistral-ai-context_window |
| Max output | — | Unknown | — | — | — | #mistral-ai-max_output |
| Image input | ● | Vendor-claimed | 2026-01-15 | https://docs.mistral.ai/capabilities/vision/ | — | #mistral-ai-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | Speech transcription models exist; no native speech output. | #mistral-ai-audio_io |
| Open weights | ◐ | Vendor-claimed | 2026-01-15 | https://docs.mistral.ai/getting-started/models/models_overview/ | Licensing is genuinely mixed — many models are Apache 2.0, others sit under a research licence requiring a commercial agreement. Check per model. | #mistral-ai-open_weights |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | — | — | #mistral-ai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #mistral-ai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #mistral-ai-throughput_tps |
| $/M input | — | Unknown | — | — | Withdrawn. The $2/M previously here is the Mistral Large 2 rate, while the row names Mistral Large 3 — which shipped at materially lower per-token pricing. Quoting the old generation's price against the new generation's name overstates the cost of the model this row is actually about, so the cell is blank until Large 3's own rate is verified. | #mistral-ai-price_in |
| $/M output | — | Unknown | — | — | Withdrawn for the same reason as the input rate: $6/M was a Mistral Large 2 figure on a Mistral Large 3 row. | #mistral-ai-price_out |
| $/M cache read | — | Unknown | — | — | — | #mistral-ai-cached_input |
| Batch discount | 50 % | Inferred | 2026-01-15 | — | Batch inference is offered at a reduced rate on La Plateforme. | #mistral-ai-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-01-15 | https://docs.mistral.ai/capabilities/function_calling/ | — | #mistral-ai-tool_use |
| Schema output | ● | Inferred | 2026-01-15 | — | — | #mistral-ai-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | Reasoning is a separate model line rather than a per-request effort dial. | #mistral-ai-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #mistral-ai-computer_use |
| MCP support | ◐ | Inferred | 2026-01-15 | — | Supported through the agents SDK rather than as a server-side connector. | #mistral-ai-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #mistral-ai-prompt_cache_ttl |
| Continuity policy | Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away. | Inferred | 2026-01-15 | — | — | #mistral-ai-continuity_policy |
| Notice period | — | Unknown | — | — | Deprecation dates are published per model but no fixed minimum window was located. | #mistral-ai-deprecation_notice |
| Pinnable versions | ● | Vendor-claimed | 2026-01-15 | https://docs.mistral.ai/getting-started/models/models_overview/ | — | #mistral-ai-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-01-15 | — | Enterprise and on-premise deployments give full control; the hosted platform's defaults are the usual abuse-monitoring retention. | #mistral-ai-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #mistral-ai-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | EU-resident by design; AU residency only via self-deployment of the open-weight models. | #mistral-ai-au_region |
| API since | 2023 | Inferred | — | — | — | #mistral-ai-launch_year |
| Positioning | European models with open weights and deployable anywhere | Inferred | — | — | — | #mistral-ai-positioning |

**Verdict.** Rarely the capability leader, consistently the answer when the blocker is jurisdiction. EU residency, an on-premise licensing path and a real open-weight lineage make it the shortlist entry that clears legal review when the American labs cannot. Check licences per model — the Apache/research split is genuinely confusing.

**Pick it when**

- EU data-residency requirements that rule out US-hosted inference
- On-premise deployment with a commercial support contract
- Teams that want an open-weight escape hatch under the hosted API

### Qwen {#alibaba-qwen}

Alibaba's model family — huge open-weight range, closed flagship

|  |  |
| --- | --- |
| Site | https://qwen.ai |
| Docs | https://www.alibabacloud.com/help/en/model-studio |
| Company | Alibaba Cloud |
| Founded | 2009 |
| Open source | Yes (Apache-2.0 (most open models)) |
| Profile | https://toolweight.dev/tools/alibaba-qwen |
| Score (default weights) | 67.4 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Qwen3-Max | Inferred | 2026-01-15 | — | The Max tier is API-only; the open-weight Qwen3 models sit a tier below. | #alibaba-qwen-flagship_model |
| Context window | 262,144 tokens | Inferred | 2026-01-15 | — | — | #alibaba-qwen-context_window |
| Max output | — | Unknown | — | — | — | #alibaba-qwen-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | Via the Qwen-VL line rather than the flagship text model — which is precisely what this column calls partial, and how Cohere and Z.ai are graded for the same arrangement. The note and the grade now agree. | #alibaba-qwen-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | Qwen-Audio and Omni variants exist as separate models. | #alibaba-qwen-audio_io |
| Open weights | ◐ | Vendor-claimed | 2026-01-15 | https://github.com/QwenLM/Qwen3 | The most prolific open-weight release programme in the category — but the Max flagship is closed. | #alibaba-qwen-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://www.alibabacloud.com/help/en/model-studio/compatibility-of-openai-with-dashscope | — | #alibaba-qwen-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #alibaba-qwen-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #alibaba-qwen-throughput_tps |
| $/M input | $1.2 /M tok | Inferred | 2026-01-15 | — | Model Studio pricing is tiered by prompt length; this is the short-context rate and longer prompts cost materially more. | #alibaba-qwen-price_in |
| $/M output | $6 /M tok | Inferred | 2026-01-15 | — | — | #alibaba-qwen-price_out |
| $/M cache read | — | Unknown | — | — | — | #alibaba-qwen-cached_input |
| Batch discount | 50 % | Inferred | 2026-01-15 | — | Model Studio offers a discounted batch mode. | #alibaba-qwen-batch_discount |
| Tool use | ● | Inferred | 2026-01-15 | — | — | #alibaba-qwen-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | — | #alibaba-qwen-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | Thinking can be toggled on the hybrid models; no graduated levels. | #alibaba-qwen-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #alibaba-qwen-computer_use |
| MCP support | ◐ | Inferred | 2026-01-15 | — | — | #alibaba-qwen-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #alibaba-qwen-prompt_cache_ttl |
| Continuity policy | Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely. | Inferred | 2026-01-15 | — | — | #alibaba-qwen-continuity_policy |
| Notice period | — | Unknown | — | — | — | #alibaba-qwen-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | Dated snapshot IDs are published for most models. | #alibaba-qwen-pinned_snapshots |
| Zero retention | — | Unknown | — | — | — | #alibaba-qwen-zero_retention |
| Trains on your data | ◐ | Inferred | 2026-01-15 | — | — | #alibaba-qwen-trains_on_data |
| AU region | ◐ | Inferred | 2026-01-15 | — | Alibaba Cloud operates a Sydney region, but Model Studio availability there should be confirmed before relying on it. | #alibaba-qwen-au_region |
| API since | 2023 | Inferred | — | — | — | #alibaba-qwen-launch_year |
| Positioning | The widest open-weight family, plus a closed flagship tier | Inferred | — | — | — | #alibaba-qwen-positioning |

**Verdict.** The best open-weight range in the category — there is a Qwen model at nearly every size and modality, mostly Apache 2.0. The hosted Max tier is a reasonable mid-price flagship but rarely the reason to be here; most teams use the open weights through a Western host and treat Model Studio as optional.

**Pick it when**

- Finding an open-weight model at a specific size or modality
- Fine-tuning where a permissive licence and a model-size ladder both matter
- Asia-Pacific deployments already on Alibaba Cloud

### Meta Llama {#meta-llama}

Open-weight Llama models, hosted almost everywhere but Meta

|  |  |
| --- | --- |
| Site | https://www.llama.com |
| Docs | https://www.llama.com/docs |
| Company | Meta |
| Founded | 2004 |
| Open source | Yes (Llama Community License) |
| Profile | https://toolweight.dev/tools/meta-llama |
| Score (default weights) | 66.6 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Llama 4 Maverick | Inferred | 2026-01-15 | — | Meta's own hosted API has never been the main distribution channel; most consumption is via routers and clouds. | #meta-llama-flagship_model |
| Context window | 1,000,000 tokens | Vendor-claimed | 2026-01-15 | https://www.llama.com/models/llama-4/ | Advertised input window; usable context depends entirely on how your host has deployed it, and few hosts serve anything close to the full figure. | #meta-llama-context_window |
| Max output | — | Unknown | — | — | Host-dependent — there is no vendor endpoint to set one. Whatever your serving stack is configured for is the answer. | #meta-llama-max_output |
| Image input | ● | Vendor-claimed | 2026-01-15 | https://www.llama.com/models/llama-4/ | — | #meta-llama-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #meta-llama-audio_io |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://www.llama.com/llama-downloads/ | Llama Community License — permissive for most commercial use but with an acceptable-use policy and a large-scale-user clause, so not OSI open source. | #meta-llama-open_weights |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | — | True of essentially every host that serves Llama, rather than a property of Meta's own endpoint. | #meta-llama-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #meta-llama-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #meta-llama-throughput_tps |
| $/M input | — | Unknown | — | — | Meta's first-party API pricing is not the reference point; compare Groq, Together, Fireworks or Bedrock rates for the same weights. | #meta-llama-price_in |
| $/M output | — | Unknown | — | — | — | #meta-llama-price_out |
| $/M cache read | — | Unknown | — | — | — | #meta-llama-cached_input |
| Batch discount | — | Unknown | — | — | — | #meta-llama-batch_discount |
| Tool use | ◐ | Community-reported | 2026-01-15 | — | The models are trained for tool calling, but reliability at depth trails the closed frontier and varies by host implementation. | #meta-llama-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | Depends on the serving stack — most open-weight hosts add grammar-constrained decoding. | #meta-llama-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #meta-llama-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #meta-llama-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #meta-llama-mcp_support |
| Cache TTL | Host-dependent | Inferred | 2026-01-15 | — | — | #meta-llama-prompt_cache_ttl |
| Continuity policy | Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you — if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears. | Inferred | 2026-01-15 | — | — | #meta-llama-continuity_policy |
| Notice period | — | Unknown | — | — | Not applicable in the usual sense — retirement is a per-host decision, not a vendor one. | #meta-llama-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | You can pin a checkpoint hash. | #meta-llama-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-01-15 | — | Conditional on how you run it: total if you self-host, and entirely that host's policy if you do not. Regraded from yes because the weights alone guarantee nothing — the self-host baseline row is where an unconditional yes belongs. | #meta-llama-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #meta-llama-trains_on_data |
| AU region | ● | Inferred | 2026-01-15 | — | Deployable anywhere, including Bedrock ap-southeast-2 and your own Australian hardware. | #meta-llama-au_region |
| API since | 2023 | Inferred | — | — | — | #meta-llama-launch_year |
| Positioning | Open-weight models you can host anywhere, forever | Inferred | — | — | — | #meta-llama-positioning |

**Verdict.** Listed here as a model family rather than a serious first-party endpoint — nobody should be calling Meta's API when Groq, Together, Fireworks and Bedrock all serve the same weights better. Its real value is as the category's continuity insurance: a capable model that literally cannot be deprecated.

**Pick it when**

- Workloads that must survive any single vendor disappearing
- On-premise or air-gapped deployments with a recognisable licence
- Fine-tuning on your own data without a vendor in the loop

### Self-hosted (vLLM) {#self-hosted-vllm}

Run open weights on your own GPUs behind an OpenAI-shaped API

|  |  |
| --- | --- |
| Site | https://docs.vllm.ai |
| Repo | https://github.com/vllm-project/vllm |
| Company | Self-hosted (vLLM) |
| Founded | 2023 |
| Open source | Yes (Apache-2.0) |
| Profile | https://toolweight.dev/tools/self-hosted-vllm |
| Score (default weights) | 65.1 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Whatever you deploy — DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4 | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-flagship_model |
| Context window | — | Unknown | — | — | Set by the checkpoint and by how much KV cache your VRAM budget allows. | #self-hosted-vllm-context_window |
| Max output | — | Unknown | — | — | Whatever you configure the server for, bounded by the same VRAM budget as the input window. | #self-hosted-vllm-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-audio_io |
| Open weights | ● | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html | vLLM serves /v1/chat/completions natively, so client code written against OpenAI works unchanged. | #self-hosted-vllm-openai_compatible |
| TTFT p50 | — | Unknown | — | — | Entirely determined by your hardware, batch size and prefix-cache hit rate. | #self-hosted-vllm-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #self-hosted-vllm-throughput_tps |
| $/M input | — | Unknown | — | — | You pay GPU-hours, not tokens. An 8xH100 node runs roughly US$20–30/hour on demand, so the per-token cost depends entirely on utilisation — cheap at high load, ruinous at low. | #self-hosted-vllm-price_in |
| $/M output | — | Unknown | — | — | — | #self-hosted-vllm-price_out |
| $/M cache read | — | Unknown | — | — | Prefix caching is free and lives in VRAM; there is no separate cache-read rate. | #self-hosted-vllm-cached_input |
| Batch discount | 0 % | Inferred | — | — | Not applicable — continuous batching is how the server works, not a pricing tier. | #self-hosted-vllm-batch_discount |
| Tool use | ◐ | Inferred | 2026-01-15 | — | Tool-call parsing depends on a per-model parser plugin and is noticeably more fragile than a hosted API. | #self-hosted-vllm-tool_use |
| Schema output | ● | Vendor-claimed | 2026-01-15 | https://docs.vllm.ai/en/latest/features/structured_outputs.html | Guided decoding with JSON Schema, regex or grammar backends. | #self-hosted-vllm-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-mcp_support |
| Cache TTL | Automatic prefix cache in VRAM, no TTL | Vendor-claimed | 2026-01-15 | https://docs.vllm.ai/en/latest/features/automatic_prefix_caching.html | — | #self-hosted-vllm-prompt_cache_ttl |
| Continuity policy | Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster. | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-continuity_policy |
| Notice period | — | Unknown | — | — | Not applicable — there is no vendor to give notice. | #self-hosted-vllm-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | Pin the checkpoint hash and it is byte-identical forever. | #self-hosted-vllm-pinned_snapshots |
| Zero retention | ● | Inferred | 2026-01-15 | — | The one unconditional yes in this column: there is no vendor to retain anything, and no abuse-monitoring log you did not write yourself. | #self-hosted-vllm-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #self-hosted-vllm-trains_on_data |
| AU region | ● | Inferred | 2026-01-15 | — | The only unambiguous answer to an AU-residency requirement. | #self-hosted-vllm-au_region |
| API since | 2023 | Inferred | — | — | — | #self-hosted-vllm-launch_year |
| Positioning | Your weights, your GPUs, your OpenAI-compatible endpoint | Inferred | — | — | — | #self-hosted-vllm-positioning |

**Verdict.** The reference point rather than a recommendation. Self-hosting only beats a hosted API on cost at genuinely high, steady utilisation, and it costs you the agentic features — effort control, computer use, MCP — that make the frontier APIs worth their price. Choose it for sovereignty or permanence, not to save money.

**Pick it when**

- Data that legally cannot leave your infrastructure
- Sustained high-utilisation inference where GPU-hours beat per-token billing
- Guaranteeing a model version survives indefinitely

### Together AI {#together-ai}

Serverless and dedicated hosting for open-weight models

|  |  |
| --- | --- |
| Site | https://www.together.ai |
| Docs | https://docs.together.ai |
| Company | Together AI |
| Founded | 2022 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/together-ai |
| Score (default weights) | 64.2 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Open-weight catalogue — DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4 | Vendor-claimed | 2026-01-15 | https://docs.together.ai/docs/serverless-models | Priced per model by size class rather than as a single flagship rate. | #together-ai-flagship_model |
| Context window | — | Unknown | — | — | Per-model. | #together-ai-context_window |
| Max output | — | Unknown | — | — | Per-model. | #together-ai-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | — | #together-ai-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | — | #together-ai-audio_io |
| Open weights | ● | Inferred | 2026-01-15 | — | Serves only open-weight models, so anything you build on here can be lifted onto your own hardware. | #together-ai-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://docs.together.ai/docs/openai-api-compatibility | — | #together-ai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #together-ai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #together-ai-throughput_tps |
| $/M input | — | Unknown | — | — | Per-model tiers roughly tracking parameter count; sub-$1/M for 70B-class models, low single digits for the largest MoE models. | #together-ai-price_in |
| $/M output | — | Unknown | — | — | — | #together-ai-price_out |
| $/M cache read | — | Unknown | — | — | — | #together-ai-cached_input |
| Batch discount | 50 % | Inferred | 2026-01-15 | — | A discounted batch endpoint is offered for supported models. | #together-ai-batch_discount |
| Tool use | ◐ | Inferred | 2026-01-15 | — | — | #together-ai-tool_use |
| Schema output | ● | Inferred | 2026-01-15 | — | JSON-schema constrained decoding is supported on most served models. | #together-ai-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #together-ai-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #together-ai-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #together-ai-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #together-ai-prompt_cache_ttl |
| Continuity policy | Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint. | Inferred | 2026-01-15 | — | — | #together-ai-continuity_policy |
| Notice period | — | Unknown | — | — | Deprecations are published but the notice window is short and not contractually fixed. | #together-ai-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | — | #together-ai-pinned_snapshots |
| Zero retention | — | Unknown | — | — | Prompts are widely understood not to be retained on the paid API by default, and dedicated endpoints isolate further — but this column's bar is exclusion from all storage including abuse-monitoring logs, and no contractual zero-retention programme was located to support a yes. Regraded from an unevidenced yes: the honest answer is that we have not checked, not that the answer is favourable. | #together-ai-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #together-ai-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #together-ai-au_region |
| API since | 2023 | Inferred | — | — | — | #together-ai-launch_year |
| Positioning | Open-weight models with a Western contract and dedicated capacity | Inferred | — | — | — | #together-ai-positioning |

**Verdict.** The pragmatic way to use Chinese open-weight models without sending data to a Chinese endpoint: same weights, US infrastructure, a contract your legal team recognises. The dedicated-endpoint path is the real differentiator once you have steady load — predictable latency and no noisy neighbours.

**Pick it when**

- Running DeepSeek or Kimi weights under a Western contract
- Steady high-volume load that justifies a dedicated GPU endpoint
- Fine-tuning open-weight models without owning hardware

### xAI {#xai}

Grok models with an OpenAI-shaped API and live X data access

|  |  |
| --- | --- |
| Site | https://x.ai |
| Docs | https://docs.x.ai |
| Company | xAI |
| Founded | 2023 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/xai |
| Score (default weights) | 63.6 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Grok 4.1 | Inferred | 2026-01-15 | — | xAI ships fast and renames tiers often; confirm the current model list before quoting. | #xai-flagship_model |
| Context window | — | Unknown | — | — | Withdrawn. The 256K figure previously carried here describes Grok 4, not the Grok 4.1 generation named above, and the 4.1 tier that most API traffic actually reaches is documented with a far larger window. Rather than publish a number that measures a different model from the one this row names, this cell is blank until we can verify the flagship's own figure. | #xai-context_window |
| Max output | — | Unknown | — | — | — | #xai-max_output |
| Image input | ● | Inferred | 2026-01-15 | — | — | #xai-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #xai-audio_io |
| Open weights | ◐ | Community-reported | 2026-01-15 | — | Older Grok generations have been released publicly; current models are closed. | #xai-open_weights |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | — | Deliberately shaped as a drop-in for the OpenAI SDKs. | #xai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #xai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #xai-throughput_tps |
| $/M input | — | Unknown | — | — | Withdrawn, and worth explaining. This cell previously read $3/M — a Grok 4-class list rate, as its own note conceded — while the row names Grok 4.1. Those are different models an order of magnitude apart in price: the 4.1 generation reaches the API mainly through a fast variant billed in cents, not dollars, per million tokens. Publishing the old number under the new model's name would have overstated the cost of running Grok 4.1 by more than tenfold, which is exactly the failure this page's methodology promises not to make. It stays blank until we can quote a rate for the model actually named. | #xai-price_in |
| $/M output | — | Unknown | — | — | Withdrawn for the same reason as the input rate: the $15/M previously here was a Grok 4 figure standing in for a Grok 4.1 row. | #xai-price_out |
| $/M cache read | — | Unknown | — | — | — | #xai-cached_input |
| Batch discount | — | Unknown | — | — | No batch endpoint documented at the time of writing. | #xai-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-01-15 | https://docs.x.ai/docs/guides/function-calling | — | #xai-tool_use |
| Schema output | ● | Inferred | 2026-01-15 | — | — | #xai-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | Reasoning is exposed on specific model variants rather than as a graduated per-request control across the line-up. | #xai-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #xai-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #xai-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #xai-prompt_cache_ttl |
| Continuity policy | Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here. | Inferred | 2026-01-15 | — | — | #xai-continuity_policy |
| Notice period | — | Unknown | — | — | No published policy located. | #xai-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | — | #xai-pinned_snapshots |
| Zero retention | — | Unknown | — | — | — | #xai-zero_retention |
| Trains on your data | ◐ | Inferred | 2026-01-15 | — | Consumer terms are permissive; API terms are stricter but read them rather than assuming parity with OpenAI or Anthropic. | #xai-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #xai-au_region |
| API since | 2024 | Inferred | — | — | — | #xai-launch_year |
| Positioning | Fast, cheap frontier models with live access to X | Inferred | — | — | — | #xai-positioning |

**Verdict.** Trivially easy to try given the OpenAI-shaped API, and the fast tiers are widely reported as some of the best price-performance in the category — but note that this row currently carries no price or context figure at all. The numbers that were here described Grok 4 while the row named Grok 4.1, and a wrong number is worse than a blank one, so they were pulled rather than patched. Read the pricing page yourself before budgeting. The governance story is separately the thinnest of the first-party labs — no published deprecation policy, no regional pinning, and terms that need reading rather than assuming. Fine for consumer features, harder to justify for regulated work.

**Pick it when**

- Products that want real-time X/Twitter context in the model
- Cheap evaluation of a frontier alternative with a one-line base-URL change
- Consumer applications where governance requirements are light

### Fireworks AI {#fireworks-ai}

Fast open-weight inference with strong structured-output support

|  |  |
| --- | --- |
| Site | https://fireworks.ai |
| Docs | https://docs.fireworks.ai |
| Company | Fireworks AI |
| Founded | 2022 |
| Funding | Venture-backed |
| Open source | No |
| Profile | https://toolweight.dev/tools/fireworks-ai |
| Score (default weights) | 63.2 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Open-weight catalogue — DeepSeek, Kimi K2, Qwen3, Llama 4 | Vendor-claimed | 2026-01-15 | https://docs.fireworks.ai/models/overview | — | #fireworks-ai-flagship_model |
| Context window | — | Unknown | — | — | Per-model. | #fireworks-ai-context_window |
| Max output | — | Unknown | — | — | Per-model. | #fireworks-ai-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | — | #fireworks-ai-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | Transcription models are offered; no native speech output. | #fireworks-ai-audio_io |
| Open weights | ● | Inferred | 2026-01-15 | — | — | #fireworks-ai-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://docs.fireworks.ai/tools-sdks/openai-compatibility | — | #fireworks-ai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #fireworks-ai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #fireworks-ai-throughput_tps |
| $/M input | — | Unknown | — | — | Per-model tiers by parameter count, comparable to Together. | #fireworks-ai-price_in |
| $/M output | — | Unknown | — | — | — | #fireworks-ai-price_out |
| $/M cache read | — | Unknown | — | — | — | #fireworks-ai-cached_input |
| Batch discount | — | Unknown | — | — | — | #fireworks-ai-batch_discount |
| Tool use | ◐ | Inferred | 2026-01-15 | — | — | #fireworks-ai-tool_use |
| Schema output | ● | Vendor-claimed | 2026-01-15 | https://docs.fireworks.ai/structured-responses/structured-response-formatting | Grammar-based and JSON-schema constrained decoding, more reliable than most open-weight hosts. | #fireworks-ai-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #fireworks-ai-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #fireworks-ai-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #fireworks-ai-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #fireworks-ai-prompt_cache_ttl |
| Continuity policy | Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient. | Inferred | 2026-01-15 | — | — | #fireworks-ai-continuity_policy |
| Notice period | — | Unknown | — | — | — | #fireworks-ai-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | — | #fireworks-ai-pinned_snapshots |
| Zero retention | — | Unknown | — | — | No contractual zero-retention programme was located, and this column's bar includes abuse-monitoring logs. Regraded from an unevidenced yes — absence of a published retention policy is not evidence of a favourable one. | #fireworks-ai-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #fireworks-ai-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #fireworks-ai-au_region |
| API since | 2023 | Inferred | — | — | — | #fireworks-ai-launch_year |
| Positioning | Low-latency open-weight inference with strict structured output | Inferred | — | — | — | #fireworks-ai-positioning |

**Verdict.** Hard to separate from Together on paper; the practical split is that Fireworks invests more in constrained decoding and latency tuning. If your pipeline breaks when a response fails to parse, its grammar enforcement is the more reliable of the two. Benchmark both on your own workload — the difference is real but small.

**Pick it when**

- Open-weight workloads where output must validate against a schema every time
- Latency-sensitive serving without moving to custom silicon
- Fine-tuned open models served on managed infrastructure

### Cohere {#cohere}

Enterprise-focused models built for RAG and private deployment

|  |  |
| --- | --- |
| Site | https://cohere.com |
| Docs | https://docs.cohere.com |
| Company | Cohere |
| Founded | 2019 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/cohere |
| Score (default weights) | 62.9 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Command A (command-a-03-2025) | Inferred | 2026-01-15 | — | — | #cohere-flagship_model |
| Context window | 256,000 tokens | Inferred | 2026-01-15 | — | — | #cohere-context_window |
| Max output | — | Unknown | — | — | — | #cohere-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | — | #cohere-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #cohere-audio_io |
| Open weights | ◐ | Inferred | 2026-01-15 | — | Weights are published for research under a non-commercial licence; commercial use requires a Cohere agreement. | #cohere-open_weights |
| OpenAI-compat API | ◐ | Inferred | 2026-01-15 | — | Cohere's native API is its own shape; a compatibility layer exists but is not the documented path. | #cohere-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #cohere-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #cohere-throughput_tps |
| $/M input | $2.5 /M tok | Inferred | 2026-01-15 | https://cohere.com/pricing | — | #cohere-price_in |
| $/M output | $10 /M tok | Inferred | 2026-01-15 | — | — | #cohere-price_out |
| $/M cache read | — | Unknown | — | — | — | #cohere-cached_input |
| Batch discount | — | Unknown | — | — | — | #cohere-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-01-15 | https://docs.cohere.com/docs/tool-use | — | #cohere-tool_use |
| Schema output | ● | Inferred | 2026-01-15 | — | — | #cohere-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #cohere-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #cohere-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #cohere-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #cohere-prompt_cache_ttl |
| Continuity policy | Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does. | Inferred | 2026-01-15 | — | — | #cohere-continuity_policy |
| Notice period | — | Unknown | — | — | — | #cohere-deprecation_notice |
| Pinnable versions | ● | Vendor-claimed | 2026-01-15 | https://docs.cohere.com/docs/models | — | #cohere-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-01-15 | — | Private deployment is the primary answer, and enterprise terms cover the hosted platform — both of which are 'available on request or on a higher tier', which is what this column calls partial. Regraded down from yes: a deployment mode you have to buy is not the same as retention being off by default. | #cohere-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | Not on the paid platform; the free trial tier is a different agreement. | #cohere-trains_on_data |
| AU region | ◐ | Inferred | 2026-01-15 | — | Deployable into any cloud region you control, including AU; the hosted endpoint is not AU-resident. | #cohere-au_region |
| API since | 2021 | Inferred | — | — | — | #cohere-launch_year |
| Positioning | Enterprise RAG models you can deploy inside your own network | Inferred | — | — | — | #cohere-positioning |

**Verdict.** Has sensibly stopped chasing the frontier and now competes where it can win: retrieval quality, private deployment and enterprise contracts. Its rerank and embedding models remain best-in-class and are the more common reason to be a customer. As a general-purpose chat API it is priced like a frontier lab without matching one.

**Pick it when**

- RAG pipelines where rerank quality drives the result
- Deployments that must run inside your own VPC or data centre
- Enterprises whose procurement needs a vendor agreement, not a credit card

### Anthropic {#anthropic}

Claude models, built around long agentic runs and tool use

|  |  |
| --- | --- |
| Site | https://www.anthropic.com |
| Docs | https://platform.claude.com/docs |
| Company | Anthropic |
| Founded | 2021 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/anthropic |
| Score (default weights) | 62.7 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Claude Fable 5 (claude-fable-5) | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/about-claude/models/overview | Opus 4.8 at $5/$25 is the practical default for most builds; Sonnet 5 at $3/$15 list and Haiku 4.5 at $1/$5 cover volume and latency. Two Fable 5 properties do not generalise from the rest of the family: thinking is always on and cannot be disabled, and a request can come back HTTP 200 with stop_reason 'refusal' and empty or partial content, so client code must branch on stop_reason before it reads content. A server-side fallbacks parameter re-runs the same request on another model when that happens — budget for that path rather than discovering it in production. | #anthropic-flagship_model |
| Context window | 1,000,000 tokens | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/about-claude/models/overview | Input window. Output is capped separately at 128K. | #anthropic-context_window |
| Max output | 128,000 tokens | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/about-claude/models/overview | 128K on Fable 5, Opus 4.8 and Sonnet 5; Haiku 4.5 caps at 64K. Streaming is required in practice for anything near the ceiling. | #anthropic-max_output |
| Image input | ● | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/build-with-claude/vision | High-resolution input up to 2576px on the long edge since Opus 4.7; coordinates map 1:1 to image pixels. | #anthropic-vision |
| Audio in/out | ○ | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs | No speech input or output on the Messages API. Pair with a separate ASR/TTS vendor. | #anthropic-audio_io |
| Open weights | ○ | Inferred | 2026-06-24 | — | — | #anthropic-open_weights |
| OpenAI-compat API | ◐ | Inferred | 2026-06-24 | — | A chat-completions compatibility layer exists for porting, but tool use, thinking blocks, effort and caching all require the native API. | #anthropic-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #anthropic-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #anthropic-throughput_tps |
| $/M input | $10 /M tok | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/pricing | Fable 5 rate. Opus 4.8 is $5/M, Sonnet 5 $3/M list, Haiku 4.5 $1/M. Sonnet 5 carries an introductory $2/M through 2026-08-31, which is today's actual rate. | #anthropic-price_in |
| $/M output | $50 /M tok | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/pricing | Fable 5 rate. Opus 4.8 is $25/M, Sonnet 5 $15/M list, Haiku 4.5 $5/M. Sonnet 5 carries an introductory $10/M through 2026-08-31, which is today's actual rate. | #anthropic-price_out |
| $/M cache read | $1 /M tok | Inferred | 2026-06-24 | https://platform.claude.com/docs/en/build-with-claude/prompt-caching | Cache reads bill at roughly 0.1x input. Writes cost 1.25x at the 5-minute TTL, where two requests break even (1.25x + 0.1x against 2x uncached), and 2x at the 1-hour TTL, where you need three. The longer TTL survives gaps in bursty traffic; the doubled write is what it costs you. | #anthropic-cached_input |
| Batch discount | 50 % | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/build-with-claude/batch-processing | — | #anthropic-batch_discount |
| Tool use | ● | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview | Parallel calls on by default; SDK tool runners drive the loop with per-turn approval hooks. | #anthropic-tool_use |
| Schema output | ● | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/build-with-claude/structured-outputs | output_config.format enforces a JSON Schema; strict:true does the same for tool parameters. Recursive schemas and numeric bounds are unsupported. | #anthropic-structured_output |
| Effort control | ● | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/build-with-claude/effort | Five levels — low, medium, high, xhigh, max — plus adaptive thinking. Fixed token budgets were removed on the current generation. | #anthropic-reasoning_effort |
| Computer use | ● | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use | Graded yes because the flagship drives it directly and returns coordinates itself. Still a beta, and it needs a beta header — but so does every implementation in this column, so beta status is not what separates yes from partial here. Self-hosted or Anthropic-hosted environments. | #anthropic-computer_use |
| MCP support | ● | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/agents-and-tools/mcp-connector | Anthropic authored MCP. Server connections are available on the Messages API and first-class in Managed Agents. | #anthropic-mcp_support |
| Cache TTL | 5 min default, 1 h option; automatic or explicit breakpoints | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/build-with-claude/prompt-caching | Two paths: a top-level automatic setting that caches the last cacheable block, or up to 4 explicit breakpoints when you need to choose the boundary yourself. The minimum cacheable prefix is model-dependent and silently applied — 2048 tokens on the Fable 5 flagship, 4096 on the Opus tier and Haiku 4.5. A prefix under the threshold does not error; it simply never caches. | #anthropic-prompt_cache_ttl |
| Continuity policy | Published deprecation page lists a retirement date per model, typically months ahead — Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee. | Inferred | 2026-06-24 | https://platform.claude.com/docs/en/about-claude/model-deprecations | The retirement dates and the alias-only model IDs come straight from Anthropic's deprecation and models pages. The closing sentence — that the date is therefore your only guarantee — is toolweight's reading, not a vendor statement, which is why this cell is graded inferred rather than vendor-claimed. | #anthropic-continuity_policy |
| Notice period | 180 days | Inferred | — | — | Observed from published retirement dates rather than a contractual guarantee; recent retirements have run six months or more from announcement. | #anthropic-deprecation_notice |
| Pinnable versions | ◐ | Vendor-claimed | 2026-06-24 | https://platform.claude.com/docs/en/about-claude/models/overview | Haiku 4.5 and older models expose dated IDs. Fable 5, Opus 4.8 and Sonnet 5 are alias-only. | #anthropic-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-06-24 | — | Zero data retention is available to eligible organisations, but Fable 5 requires 30-day retention and rejects requests from ZDR orgs outright. | #anthropic-zero_retention |
| Trains on your data | ○ | Inferred | 2026-06-24 | — | API inputs and outputs are not used to train models by default. | #anthropic-trains_on_data |
| AU region | ◐ | Inferred | 2026-06-24 | — | First-party inference_geo covers US and EU residency. AU-resident serving is via Bedrock ap-southeast-2 or Vertex australia-southeast1, not the first-party endpoint. | #anthropic-au_region |
| API since | 2023 | Inferred | — | — | — | #anthropic-launch_year |
| Positioning | Frontier models for long-horizon agentic work and code | Inferred | — | — | — | #anthropic-positioning |

**Verdict.** The best platform here for anything that runs a tool loop for more than a few turns — effort control, explicit cache breakpoints and MCP are designed for that shape of work rather than retrofitted. The continuity story regressed with the current generation: dropping dated snapshot IDs means you pin to a published retirement date, not to an immutable model. One flagship-specific trap worth wiring for on day one: Fable 5 can decline a request outright, returning HTTP 200 with stop_reason 'refusal' and no usable content, so a client that reads the first content block unconditionally breaks rather than errors. Anthropic ships a fallbacks parameter that re-runs the request on another model; use it, or handle the stop reason yourself.

**Pick it when**

- Agent harnesses that call tools for minutes at a time
- Code generation and review where first-try correctness pays for the token price
- Teams that want cost tuned per route via graduated effort levels

### OpenRouter {#openrouter}

One OpenAI-shaped key in front of hundreds of models

|  |  |
| --- | --- |
| Site | https://openrouter.ai |
| Docs | https://openrouter.ai/docs |
| Company | OpenRouter |
| Founded | 2023 |
| Open source | No |
| Profile | https://toolweight.dev/tools/openrouter |
| Score (default weights) | 60.9 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Whichever upstream model you route to (400+ available) | Vendor-claimed | 2026-01-15 | https://openrouter.ai/models | — | #openrouter-flagship_model |
| Context window | — | Unknown | — | — | Per-model; the router exposes each upstream's window unchanged. | #openrouter-context_window |
| Max output | — | Unknown | — | — | Per-model, passed through from the upstream. | #openrouter-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | Available on models that support it. | #openrouter-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | — | #openrouter-audio_io |
| Open weights | ◐ | Inferred | 2026-01-15 | — | Routes to both open and closed models; publishes none. | #openrouter-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://openrouter.ai/docs/api-reference/overview | A single OpenAI-shaped endpoint in front of every upstream is the entire product. | #openrouter-openai_compatible |
| TTFT p50 | — | Unknown | — | — | OpenRouter publishes live per-model latency and throughput in its own console — better data than this column can hold. | #openrouter-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #openrouter-throughput_tps |
| $/M input | — | Unknown | — | — | Upstream list price passed through, plus a fee on credit purchases (around 5%). No inference markup. | #openrouter-price_in |
| $/M output | — | Unknown | — | — | — | #openrouter-price_out |
| $/M cache read | — | Unknown | — | — | — | #openrouter-cached_input |
| Batch discount | 0 % | Inferred | 2026-01-15 | — | No batch endpoint — a real cost if half your workload could tolerate async. | #openrouter-batch_discount |
| Tool use | ● | Inferred | 2026-01-15 | — | Passed through where the upstream supports it. | #openrouter-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | — | #openrouter-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | A normalised reasoning parameter is mapped onto each upstream's equivalent where one exists. | #openrouter-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #openrouter-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #openrouter-mcp_support |
| Cache TTL | Passed through where the upstream supports caching | Inferred | 2026-01-15 | — | — | #openrouter-prompt_cache_ttl |
| Continuity policy | Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself — a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement. | Inferred | 2026-01-15 | — | — | #openrouter-continuity_policy |
| Notice period | — | Unknown | — | — | Inherits each upstream's policy; OpenRouter publishes none of its own. | #openrouter-deprecation_notice |
| Pinnable versions | ◐ | Inferred | 2026-01-15 | — | Model IDs are stable but resolve to whichever upstream host is available, so serving behaviour can vary between identical requests. | #openrouter-pinned_snapshots |
| Zero retention | ◐ | Inferred | 2026-01-15 | — | Logging is off by default at the router and provider policies can be filtered so requests never reach a host that trains on prompts — but the router cannot make a promise on behalf of an upstream's abuse-monitoring logs, which is the bar this column sets. Regraded down from yes: you are buying a filter over other people's retention terms, not zero retention. | #openrouter-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #openrouter-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #openrouter-au_region |
| API since | 2023 | Inferred | — | — | — | #openrouter-launch_year |
| Positioning | One key, one API shape, every model worth calling | Inferred | — | — | — | #openrouter-positioning |

**Verdict.** The highest-leverage integration in the category: one endpoint, hundreds of models, automatic failover, and privacy filters that let you exclude hosts who train on prompts. The trade is that you get the intersection of upstream features, not the union — no batch, no computer use, and caching only where it passes through.

**Pick it when**

- Keeping model choice a runtime decision rather than an architectural one
- Benchmarking a dozen models without a dozen accounts
- Failover when a single upstream host goes down

### Z.ai (GLM) {#zhipu-zai}

GLM models with an unusually cheap flat-rate coding plan

|  |  |
| --- | --- |
| Site | https://z.ai |
| Docs | https://docs.z.ai |
| Company | Zhipu AI |
| Founded | 2019 |
| Open source | Yes (MIT) |
| Profile | https://toolweight.dev/tools/zhipu-zai |
| Score (default weights) | 59.9 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | GLM-4.6 | Inferred | 2026-01-15 | — | A newer GLM generation may have shipped since; verify before quoting. | #zhipu-zai-flagship_model |
| Context window | 200,000 tokens | Inferred | 2026-01-15 | — | — | #zhipu-zai-context_window |
| Max output | — | Unknown | — | — | — | #zhipu-zai-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | Via the separate GLM-V line. | #zhipu-zai-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #zhipu-zai-audio_io |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://huggingface.co/zai-org/GLM-4.6 | GLM-4.6 weights released under MIT. | #zhipu-zai-open_weights |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | — | — | #zhipu-zai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #zhipu-zai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #zhipu-zai-throughput_tps |
| $/M input | $0.6 /M tok | Inferred | 2026-01-15 | — | — | #zhipu-zai-price_in |
| $/M output | $2.2 /M tok | Inferred | 2026-01-15 | — | Token pricing is beside the point for heavy users — the flat-rate coding subscription is roughly the price of a coffee a month and covers a very large quota. | #zhipu-zai-price_out |
| $/M cache read | — | Unknown | — | — | — | #zhipu-zai-cached_input |
| Batch discount | — | Unknown | — | — | — | #zhipu-zai-batch_discount |
| Tool use | ● | Community-reported | 2026-01-15 | — | — | #zhipu-zai-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | — | #zhipu-zai-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | — | #zhipu-zai-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #zhipu-zai-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #zhipu-zai-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #zhipu-zai-prompt_cache_ttl |
| Continuity policy | Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model. | Inferred | 2026-01-15 | — | — | #zhipu-zai-continuity_policy |
| Notice period | — | Unknown | — | — | — | #zhipu-zai-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | — | #zhipu-zai-pinned_snapshots |
| Zero retention | ○ | Inferred | 2026-01-15 | — | — | #zhipu-zai-zero_retention |
| Trains on your data | ◐ | Inferred | 2026-01-15 | — | — | #zhipu-zai-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #zhipu-zai-au_region |
| API since | 2023 | Inferred | — | — | — | #zhipu-zai-launch_year |
| Positioning | Coding-focused GLM models with a flat-rate subscription | Inferred | — | — | — | #zhipu-zai-positioning |

**Verdict.** The flat-rate coding plan is the genuinely disruptive product here — it makes running an agent hard all day a fixed cost rather than a variable one, which no per-token vendor can match. Model quality is a step below the frontier on hard reasoning; for routine coding turns most users do not notice.

**Pick it when**

- Driving a coding agent all day on a predictable monthly cost
- Cheap high-volume code completion and refactoring
- Self-hosting an MIT-licensed coding model

### Moonshot AI {#moonshot-ai}

Kimi models — open-weight agentic performance at low cost

|  |  |
| --- | --- |
| Site | https://www.moonshot.ai |
| Docs | https://platform.moonshot.ai/docs |
| Company | Moonshot AI |
| Founded | 2023 |
| Open source | Yes (Modified MIT) |
| Profile | https://toolweight.dev/tools/moonshot-ai |
| Score (default weights) | 58.8 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Kimi K2 Thinking | Inferred | 2026-01-15 | — | — | #moonshot-ai-flagship_model |
| Context window | 256,000 tokens | Inferred | 2026-01-15 | — | — | #moonshot-ai-context_window |
| Max output | — | Unknown | — | — | — | #moonshot-ai-max_output |
| Image input | ○ | Inferred | 2026-01-15 | — | — | #moonshot-ai-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #moonshot-ai-audio_io |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://github.com/MoonshotAI/Kimi-K2 | Modified MIT licence with an attribution condition for very large deployments. | #moonshot-ai-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://platform.moonshot.ai/docs/api/chat | — | #moonshot-ai-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #moonshot-ai-ttft_p50 |
| Output tok/s | — | Unknown | — | — | Slow on the first-party endpoint by most community reports; Groq and Cerebras serve the same weights far faster. | #moonshot-ai-throughput_tps |
| $/M input | $0.6 /M tok | Inferred | 2026-01-15 | — | Cache-miss rate; cache hits are roughly a quarter of this. | #moonshot-ai-price_in |
| $/M output | $2.5 /M tok | Inferred | 2026-01-15 | — | — | #moonshot-ai-price_out |
| $/M cache read | $0.15 /M tok | Inferred | 2026-01-15 | — | — | #moonshot-ai-cached_input |
| Batch discount | — | Unknown | — | — | — | #moonshot-ai-batch_discount |
| Tool use | ● | Community-reported | 2026-01-15 | — | The open-weight model most often reported as competitive with closed frontier models on agentic benchmarks. | #moonshot-ai-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | — | #moonshot-ai-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | — | #moonshot-ai-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #moonshot-ai-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #moonshot-ai-mcp_support |
| Cache TTL | Explicit context caching, paid per storage hour | Inferred | 2026-01-15 | — | — | #moonshot-ai-prompt_cache_ttl |
| Continuity policy | Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence — download the checkpoint and continuity is your problem, not theirs. | Inferred | 2026-01-15 | — | — | #moonshot-ai-continuity_policy |
| Notice period | — | Unknown | — | — | — | #moonshot-ai-deprecation_notice |
| Pinnable versions | ● | Inferred | 2026-01-15 | — | — | #moonshot-ai-pinned_snapshots |
| Zero retention | ○ | Inferred | 2026-01-15 | — | — | #moonshot-ai-zero_retention |
| Trains on your data | ◐ | Inferred | 2026-01-15 | — | — | #moonshot-ai-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #moonshot-ai-au_region |
| API since | 2023 | Inferred | — | — | — | #moonshot-ai-launch_year |
| Positioning | Open-weight agentic performance at open-weight prices | Inferred | — | — | — | #moonshot-ai-positioning |

**Verdict.** Kimi K2 is the strongest argument that open weights have caught up on agentic work specifically — it holds up in tool loops where other open models fall apart. Reach it through Groq, Fireworks or Together rather than the first-party endpoint, which is slower and has weaker data terms.

**Pick it when**

- Agent loops on an open-weight budget
- Self-hosting a tool-use-capable model under a permissive licence
- Replacing a frontier model on the cheaper half of a mixed workload

### Groq {#groq}

Custom LPU silicon serving open-weight models at extreme speed

|  |  |
| --- | --- |
| Site | https://groq.com |
| Docs | https://console.groq.com/docs |
| Company | Groq |
| Founded | 2016 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/groq |
| Score (default weights) | 56.6 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Curated open-weight models (Kimi K2, Llama 4, Qwen3, GPT-OSS) | Vendor-claimed | 2026-01-15 | https://console.groq.com/docs/models | The menu is short and rotates; check the models page before designing around one. | #groq-flagship_model |
| Context window | — | Unknown | — | — | Per-model. | #groq-context_window |
| Max output | — | Unknown | — | — | Per-model. | #groq-max_output |
| Image input | ◐ | Inferred | 2026-01-15 | — | — | #groq-vision |
| Audio in/out | ◐ | Inferred | 2026-01-15 | — | Whisper-class transcription is served fast; no native speech output. | #groq-audio_io |
| Open weights | ● | Inferred | 2026-01-15 | — | — | #groq-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://console.groq.com/docs/openai | — | #groq-openai_compatible |
| TTFT p50 | 250 ms | Community-reported | 2026-01-15 | — | Community-measured on short prompts from third-party benchmarks; varies by model and region. Not measured by toolweight. | #groq-ttft_p50 |
| Output tok/s | 400 tok/s | Community-reported | 2026-01-15 | — | Order-of-magnitude figure for mid-sized models — roughly 5–10x a general-purpose GPU host. Smaller models run considerably faster. | #groq-throughput_tps |
| $/M input | — | Unknown | — | — | Per-model and consistently below GPU hosts for the same weights. | #groq-price_in |
| $/M output | — | Unknown | — | — | — | #groq-price_out |
| $/M cache read | — | Unknown | — | — | — | #groq-cached_input |
| Batch discount | — | Unknown | — | — | A batch API is offered at a discount, which is unusual among the speed-focused hosts — but the percentage is disputed. The 50% previously carried here is the category norm rather than a figure confirmed for Groq, and a materially lower rate has been reported. Since this is the only pricing-family number any router or host on this page carries, it is blank rather than assumed: verify against the console before modelling a saving on it. | #groq-batch_discount |
| Tool use | ◐ | Inferred | 2026-01-15 | — | — | #groq-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | — | #groq-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #groq-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #groq-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #groq-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #groq-prompt_cache_ttl |
| Continuity policy | The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun. | Community-reported | 2026-01-15 | — | — | #groq-continuity_policy |
| Notice period | — | Unknown | — | — | Deprecations are announced but the window is short and not published as a guaranteed minimum. | #groq-deprecation_notice |
| Pinnable versions | ◐ | Inferred | 2026-01-15 | — | You can name a specific model, but it only exists while Groq chooses to host it. | #groq-pinned_snapshots |
| Zero retention | — | Unknown | — | — | No contractual zero-retention programme was located, and this column's bar includes abuse-monitoring logs. Regraded from an unevidenced yes. | #groq-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #groq-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #groq-au_region |
| API since | 2024 | Inferred | — | — | — | #groq-launch_year |
| Positioning | Custom LPU silicon for open-weight models at very low latency | Inferred | — | — | — | #groq-positioning |

**Verdict.** When the response appearing instantly is the product, Groq changes what you can build — voice, live search and inline completion feel different at these token rates. It is not a general-purpose platform: the menu is short, models rotate off with little warning, and there is no path to bring your own weights.

**Pick it when**

- Voice agents and any interface where perceived latency is the feature
- High-volume cheap inference on well-known open-weight models
- Fast transcription alongside text generation on one key

### Cerebras {#cerebras}

Wafer-scale inference — the fastest tokens per second available

|  |  |
| --- | --- |
| Site | https://cerebras.ai |
| Docs | https://inference-docs.cerebras.ai |
| Company | Cerebras |
| Founded | 2016 |
| Funding | Late-stage private |
| Open source | No |
| Profile | https://toolweight.dev/tools/cerebras |
| Score (default weights) | 56.5 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | Curated open-weight models (Qwen3, GLM, Llama, GPT-OSS) | Vendor-claimed | 2026-01-15 | https://inference-docs.cerebras.ai/models/overview | — | #cerebras-flagship_model |
| Context window | — | Unknown | — | — | Per-model. | #cerebras-context_window |
| Max output | — | Unknown | — | — | Per-model. | #cerebras-max_output |
| Image input | ○ | Inferred | 2026-01-15 | — | — | #cerebras-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #cerebras-audio_io |
| Open weights | ● | Inferred | 2026-01-15 | — | — | #cerebras-open_weights |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://inference-docs.cerebras.ai/resources/openai | — | #cerebras-openai_compatible |
| TTFT p50 | 200 ms | Community-reported | 2026-01-15 | — | Community-measured on short prompts; not measured by toolweight. | #cerebras-ttft_p50 |
| Output tok/s | 2,000 tok/s | Community-reported | 2026-01-15 | — | Order-of-magnitude figure — third-party benchmarks routinely report 1,500–3,000 tok/s on large MoE models where GPU hosts manage a few hundred. | #cerebras-throughput_tps |
| $/M input | — | Unknown | — | — | Per-model. | #cerebras-price_in |
| $/M output | — | Unknown | — | — | — | #cerebras-price_out |
| $/M cache read | — | Unknown | — | — | — | #cerebras-cached_input |
| Batch discount | — | Unknown | — | — | — | #cerebras-batch_discount |
| Tool use | ◐ | Inferred | 2026-01-15 | — | — | #cerebras-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | — | #cerebras-structured_output |
| Effort control | ○ | Inferred | 2026-01-15 | — | — | #cerebras-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #cerebras-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #cerebras-mcp_support |
| Cache TTL | — | Unknown | — | — | — | #cerebras-prompt_cache_ttl |
| Continuity policy | Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite. | Community-reported | 2026-01-15 | — | — | #cerebras-continuity_policy |
| Notice period | — | Unknown | — | — | — | #cerebras-deprecation_notice |
| Pinnable versions | ◐ | Inferred | 2026-01-15 | — | — | #cerebras-pinned_snapshots |
| Zero retention | — | Unknown | — | — | No contractual zero-retention programme was located, and this column's bar includes abuse-monitoring logs. Regraded from an unevidenced yes. | #cerebras-zero_retention |
| Trains on your data | ○ | Inferred | 2026-01-15 | — | — | #cerebras-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #cerebras-au_region |
| API since | 2024 | Inferred | — | — | — | #cerebras-launch_year |
| Positioning | Wafer-scale inference — the highest tokens per second available | Inferred | — | — | — | #cerebras-positioning |

**Verdict.** Consistently the fastest single-stream generation you can buy, by a margin that is qualitative rather than incremental — reasoning models that take half a minute elsewhere return in a couple of seconds. The catch is the same as Groq's: a short menu, no bring-your-own-weights, and no continuity guarantees at all.

**Pick it when**

- Reasoning models where thinking-token latency is the bottleneck
- Interactive tools that feel broken at 50 tokens per second
- Benchmarking how much of your UX problem is actually latency

### DeepSeek {#deepseek}

Frontier-adjacent models at a small fraction of Western prices

|  |  |
| --- | --- |
| Site | https://www.deepseek.com |
| Docs | https://api-docs.deepseek.com |
| Company | DeepSeek |
| Founded | 2023 |
| Open source | Yes (MIT) |
| Profile | https://toolweight.dev/tools/deepseek |
| Score (default weights) | 46.2 / 100 |

| Field | Value | Confidence | Verified | Source | Note | Anchor |
| --- | --- | --- | --- | --- | --- | --- |
| Flagship model | DeepSeek-V3.2 (deepseek-chat / deepseek-reasoner) | Inferred | 2026-01-15 | https://api-docs.deepseek.com/quick_start/pricing | The API exposes rolling aliases; the model behind deepseek-chat changes without a version bump. | #deepseek-flagship_model |
| Context window | 128,000 tokens | Inferred | 2026-01-15 | — | — | #deepseek-context_window |
| Max output | — | Unknown | — | — | — | #deepseek-max_output |
| Image input | ○ | Inferred | 2026-01-15 | — | — | #deepseek-vision |
| Audio in/out | ○ | Inferred | 2026-01-15 | — | — | #deepseek-audio_io |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://github.com/deepseek-ai/DeepSeek-V3 | V3 line released under MIT — among the most permissive licences of any frontier-adjacent model. | #deepseek-open_weights |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | — | — | #deepseek-openai_compatible |
| TTFT p50 | — | Unknown | — | — | — | #deepseek-ttft_p50 |
| Output tok/s | — | Unknown | — | — | — | #deepseek-throughput_tps |
| $/M input | $0.28 /M tok | Inferred | 2026-01-15 | https://api-docs.deepseek.com/quick_start/pricing | Cache-miss rate following the September 2025 price cut. | #deepseek-price_in |
| $/M output | $0.42 /M tok | Inferred | 2026-01-15 | https://api-docs.deepseek.com/quick_start/pricing | Roughly a hundred and twentieth of Claude Fable 5's $50/M, about a thirty-fifth of Claude Sonnet 5's $15/M, and about a twenty-fourth of the GPT-5-family $10/M. Which of those ratios is the honest one depends on which model you would otherwise have used — for the bulk work this rate is good for, the volume-tier comparison is the fair one. | #deepseek-price_out |
| $/M cache read | $0.028 /M tok | Inferred | 2026-01-15 | — | Cache hits are billed at a tenth of the miss rate, applied automatically. | #deepseek-cached_input |
| Batch discount | 0 % | Inferred | 2026-01-15 | — | No batch endpoint; the list price is already below most competitors' batch rates. | #deepseek-batch_discount |
| Tool use | ◐ | Community-reported | 2026-01-15 | — | Function calling works, but multi-step agentic reliability is widely reported as weaker than the closed frontier. | #deepseek-tool_use |
| Schema output | ◐ | Inferred | 2026-01-15 | — | JSON mode rather than full schema-constrained decoding. | #deepseek-structured_output |
| Effort control | ◐ | Inferred | 2026-01-15 | — | A separate reasoner model rather than a per-request effort level. | #deepseek-reasoning_effort |
| Computer use | ○ | Inferred | 2026-01-15 | — | — | #deepseek-computer_use |
| MCP support | ○ | Inferred | 2026-01-15 | — | — | #deepseek-mcp_support |
| Cache TTL | Automatic disk cache, no configuration | Inferred | 2026-01-15 | — | — | #deepseek-prompt_cache_ttl |
| Continuity policy | The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases — deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently. | Community-reported | 2026-01-15 | — | — | #deepseek-continuity_policy |
| Notice period | — | Unknown | — | — | No published deprecation policy; aliases are updated in place. | #deepseek-deprecation_notice |
| Pinnable versions | ○ | Community-reported | 2026-01-15 | — | Only rolling aliases are exposed on the first-party API. | #deepseek-pinned_snapshots |
| Zero retention | ○ | Inferred | 2026-01-15 | — | — | #deepseek-zero_retention |
| Trains on your data | ● | Community-reported | 2026-01-15 | — | Terms permit using inputs to improve services. Assume prompts are retained and read them before sending anything sensitive. | #deepseek-trains_on_data |
| AU region | ○ | Inferred | 2026-01-15 | — | — | #deepseek-au_region |
| API since | 2023 | Inferred | — | — | — | #deepseek-launch_year |
| Positioning | Near-frontier quality at a fraction of the token price | Inferred | — | — | — | #deepseek-positioning |

**Verdict.** The price floor of the category and the reason everyone else's rates fell. For bulk text work the quality-per-dollar is unmatched. Do not send regulated data to the first-party endpoint — permissive retention terms and rolling aliases with no pinning make it unsuitable for anything sensitive or long-lived. Use the MIT weights through a Western host instead.

**Pick it when**

- High-volume classification, extraction and summarisation
- Cost-sensitive workloads where a 35x output-price gap against a volume-tier frontier model outweighs a small quality gap
- Self-hosting a capable model under a genuinely permissive licence

## Methodology

Pricing on this page is stale within a fortnight. That is not hedging — this category re-prices faster than any other on toolweight, so the page carries a seven-day verification cadence and every price cell is dated. Treat any cell whose date is more than a month old as indicative and click through to the vendor's pricing page before you build a forecast on it. Figures are pay-as-you-go list rates in USD per million tokens, before batch, cache or committed-spend discounts, and before any provider-specific long-context surcharge. One convention worth stating: where a vendor is running an unexpired introductory rate, the column carries the list price and the cell note carries the introductory rate and its expiry date. An introductory rate is not a discount you negotiate — it is what you actually pay today — so read the note before modelling a bill.

Entries here are providers, not models. Because a provider ships a dozen models at a dozen prices, every price, context and capability figure on a row describes that provider's current flagship model, named explicitly in the Flagship model column so no number is orphaned from the thing it measures. Compare rows knowing that a cheaper flagship often sits beside a cheaper mid-tier that would serve your workload better. The killer column, deprecation and continuity policy, records two things: how much notice you get before a model is retired, and whether you can pin a dated snapshot that keeps serving after the alias moves on. Continuity is judged from published deprecation pages and observed retirements, not from marketing claims, and where a provider publishes no policy at all we say so rather than guessing a number. Latency figures are community-reported from third-party benchmarks; toolweight has not run its own TTFT harness yet, so most of that column is honestly blank.

## FAQ

### What actually happens when a model I depend on is deprecated?

The alias stops resolving and requests 404. If you pinned a dated snapshot you keep serving until that snapshot's published retirement date; if you called a bare alias you are migrating today. Anthropic and OpenAI publish retirement dates months ahead on dedicated deprecation pages. Several providers publish nothing, and rolling aliases can change behaviour underneath you with no version bump at all.

### Is DeepSeek or Kimi really good enough to replace Claude or GPT?

For extraction, classification, summarisation, translation and bulk rewriting, yes — and at roughly a thirty-fifth of the output-token cost of Anthropic's volume model, or about a hundred and twentieth of its top-end one. Which multiple you should care about depends on which model you would otherwise have used; the honest comparison for bulk work is against the volume tier, not the flagship. For long agentic runs with many tool calls, sustained instruction following, and code that has to compile first time, the frontier labs are still measurably ahead. Split the workload rather than picking one provider for everything.

### Does an OpenAI-compatible endpoint actually make switching easy?

It makes the transport identical and the semantics different. Chat completions port cleanly. Tool-call formats, reasoning-effort parameters, cached-token accounting, thinking blocks and structured-output enforcement do not, and prompts tuned against one model's instruction-following behaviour regress on another. Budget a re-evaluation, not a config change.

### How much does prompt caching really save on an agent loop?

More than any other lever. Cache reads typically bill at 10–25% of the input rate, and an agent loop resends the entire conversation on every turn, so a long run can be 80–90% cache reads. The trap is that caching is a prefix match: a timestamp or a reordered tool list at the front of the prompt silently invalidates everything after it.

### Which providers can serve inference from an Australian region?

Through the hyperscalers, reliably: Amazon Bedrock in ap-southeast-2 and Google Vertex AI in australia-southeast1 both serve frontier models in-country. First-party endpoints from Anthropic, OpenAI and xAI route to US or EU infrastructure by default, with data-residency controls that vary by plan. Self-hosting is the only answer that is unambiguously AU-resident.

### Should I go through a router like OpenRouter instead of direct?

Use a router when model choice is a runtime decision, when you want a fallback that survives an upstream outage, or when you are still benchmarking. Go direct when you depend on provider-specific features — prompt caching breakpoints, computer use, batch endpoints, enterprise retention terms — because routers expose the intersection of what upstreams support, not the union.

### Why is the time-to-first-token column mostly empty?

Because almost nobody publishes it, and self-reported latency is meaningless without a specified prompt length, region and concurrency. The figures shown are community measurements from third-party benchmarks for providers whose entire pitch is speed. toolweight will fill this column when it runs its own harness with published methodology, not before.

## Recent changes

| Date | Kind | Event | Tools | Source |
| --- | --- | --- | --- | --- |
| 2025-08-05 | launch | **OpenAI releases open-weight gpt-oss models** — gpt-oss-120b and gpt-oss-20b shipped under Apache 2.0, OpenAI's first open weights since GPT-2. Both landed on Groq, Together, Fireworks and Cerebras within days, and reset expectations that the frontier labs would keep everything closed. | openai, groq, cerebras, together-ai, fireworks-ai | https://openai.com/index/introducing-gpt-oss/ |
| 2025-09-29 | pricing | **DeepSeek V3.2 cuts API prices by more than half** — Sparse attention in V3.2-Exp brought long-context serving costs down and DeepSeek passed it straight through, landing input around $0.28/M and output around $0.42/M. Every budget provider on this page re-priced within the quarter. | deepseek | https://api-docs.deepseek.com/news/news250929 |
| 2025-11-18 | launch | **Google launches Gemini 3 Pro** — Gemini 3 Pro shipped across AI Studio and Vertex with a million-token window and a thinking-level control, priced well below the Opus tier for short-context work. It made long context a commodity rather than a premium feature. | google-gemini | https://blog.google/products/gemini/gemini-3/ |
| 2026-01-05 | deprecation | **Anthropic retires Claude Opus 3** — The original Opus stopped serving and its model ID began returning 404. Anthropic had published the retirement date well in advance on its deprecations page — the clearest example in the category of a lab giving developers a fixed date to plan a migration against. | anthropic | https://platform.claude.com/docs/en/about-claude/model-deprecations |
| 2026-02-19 | deprecation | **Claude Sonnet 3.7 and Haiku 3.5 retired** — Two widely deployed workhorse models went dark on the same day. Applications pinned to the dated snapshots had a hard cutover; those calling bare aliases had already been moved. A useful reminder that pinning buys you a deadline, not immunity. | anthropic | https://platform.claude.com/docs/en/about-claude/model-deprecations |

## Sources

| Source | Last verified |
| --- | --- |
| https://platform.claude.com/docs/en/about-claude/models/overview | 2026-06-24 |
| https://platform.claude.com/docs/en/build-with-claude/vision | 2026-06-24 |
| https://platform.claude.com/docs | 2026-06-24 |
| https://platform.claude.com/docs/en/pricing | 2026-06-24 |
| https://platform.claude.com/docs/en/build-with-claude/prompt-caching | 2026-06-24 |
| https://platform.claude.com/docs/en/build-with-claude/batch-processing | 2026-06-24 |
| https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview | 2026-06-24 |
| https://platform.claude.com/docs/en/build-with-claude/structured-outputs | 2026-06-24 |
| https://platform.claude.com/docs/en/build-with-claude/effort | 2026-06-24 |
| https://platform.claude.com/docs/en/agents-and-tools/tool-use/computer-use | 2026-06-24 |
| https://platform.claude.com/docs/en/agents-and-tools/mcp-connector | 2026-06-24 |
| https://platform.claude.com/docs/en/about-claude/model-deprecations | 2026-06-24 |
| https://platform.openai.com/docs/guides/images-vision | 2026-01-15 |
| https://platform.openai.com/docs/guides/realtime | 2026-01-15 |
| https://openai.com/index/introducing-gpt-oss/ | 2026-01-15 |
| https://openai.com/api/pricing/ | 2026-01-15 |
| https://platform.openai.com/docs/guides/batch | 2026-01-15 |
| https://platform.openai.com/docs/guides/function-calling | 2026-01-15 |
| https://platform.openai.com/docs/guides/structured-outputs | 2026-01-15 |
| https://platform.openai.com/docs/guides/reasoning | 2026-01-15 |
| https://platform.openai.com/docs/guides/tools-computer-use | 2026-01-15 |
| https://platform.openai.com/docs/guides/tools-remote-mcp | 2026-01-15 |
| https://platform.openai.com/docs/models | 2026-01-15 |
| https://platform.openai.com/docs/guides/your-data | 2026-01-15 |
| https://blog.google/products/gemini/gemini-3/ | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/models | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/image-understanding | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/live | 2026-01-15 |
| https://ai.google.dev/gemma | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/openai | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/pricing | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/batch-mode | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/function-calling | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/structured-output | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/thinking | 2026-01-15 |
| https://ai.google.dev/gemini-api/docs/computer-use | 2026-01-15 |
| https://docs.x.ai/docs/guides/function-calling | 2026-01-15 |
| https://www.llama.com/models/llama-4/ | 2026-01-15 |
| https://www.llama.com/llama-downloads/ | 2026-01-15 |
| https://docs.mistral.ai/capabilities/vision/ | 2026-01-15 |
| https://docs.mistral.ai/getting-started/models/models_overview/ | 2026-01-15 |
| https://docs.mistral.ai/capabilities/function_calling/ | 2026-01-15 |
| https://api-docs.deepseek.com/quick_start/pricing | 2026-01-15 |
| https://github.com/deepseek-ai/DeepSeek-V3 | 2026-01-15 |
| https://github.com/QwenLM/Qwen3 | 2026-01-15 |
| https://www.alibabacloud.com/help/en/model-studio/compatibility-of-openai-with-dashscope | 2026-01-15 |
| https://github.com/MoonshotAI/Kimi-K2 | 2026-01-15 |
| https://platform.moonshot.ai/docs/api/chat | 2026-01-15 |
| https://huggingface.co/zai-org/GLM-4.6 | 2026-01-15 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/conversation-inference-supported-models-features.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/tool-use.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/structured-output.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/computer-use.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/model-lifecycle.html | 2026-06-24 |
| https://docs.aws.amazon.com/bedrock/latest/userguide/data-protection.html | 2026-01-15 |
| https://cohere.com/pricing | 2026-01-15 |
| https://docs.cohere.com/docs/tool-use | 2026-01-15 |
| https://docs.cohere.com/docs/models | 2026-01-15 |
| https://openrouter.ai/models | 2026-01-15 |
| https://openrouter.ai/docs/api-reference/overview | 2026-01-15 |
| https://docs.together.ai/docs/serverless-models | 2026-01-15 |
| https://docs.together.ai/docs/openai-api-compatibility | 2026-01-15 |
| https://docs.fireworks.ai/models/overview | 2026-01-15 |
| https://docs.fireworks.ai/tools-sdks/openai-compatibility | 2026-01-15 |
| https://docs.fireworks.ai/structured-responses/structured-response-formatting | 2026-01-15 |
| https://console.groq.com/docs/models | 2026-01-15 |
| https://console.groq.com/docs/openai | 2026-01-15 |
| https://inference-docs.cerebras.ai/models/overview | 2026-01-15 |
| https://inference-docs.cerebras.ai/resources/openai | 2026-01-15 |
| https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html | 2026-01-15 |
| https://docs.vllm.ai/en/latest/features/structured_outputs.html | 2026-01-15 |
| https://docs.vllm.ai/en/latest/features/automatic_prefix_caching.html | 2026-01-15 |

## Related comparisons

- [Coding agents](https://toolweight.dev/compare/coding-agents) — Which AI coding agent should I use in 2026?
- [Sandbox providers](https://toolweight.dev/compare/sandbox-providers) — Which sandbox provider should I run untrusted or agent-generated code on?
- [Browser APIs](https://toolweight.dev/compare/browser-automation-apis) — Which hosted browser automation API should I use for AI agents and scraping?
- [Convex alternatives](https://toolweight.dev/compare/convex-alternatives) — What are the best alternatives to Convex, and how hard is each one to leave?

## Licence and attribution

Data from toolweight (https://toolweight.dev), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.dev/compare/llm-apis
- Machine-readable: https://toolweight.dev/compare/llm-apis.md · https://toolweight.dev/api/v1 · https://toolweight.dev/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.dev/suggest
