Anthropic, OpenAI and Google lead on capability, cheaper Chinese labs win on price, routers trade features for reach; pick on continuity, not headline price.
Providers
18
Fields compared
27
Source confidence
34%
Last verified
2026-07-30 (1mo ago)
Re-verified
every 7 days
Data confidence0% of 369 figures verified in the last 7d
Frontier LLM API providers, 18 tools compared across 27 of 27 fields.
Tool
Weighted
w5
w3
w4
w3
w4
w5
w4
w4
w8
w9
w6
w4
w8
w7
w6
w3
w5
w7
w7
w6
w7
w4
w1
Google GeminiGemini via AI Studio for prototyping or Vertex AI for production
76.9
Gemini 3.1 Pro Preview (gemini-3.1-pro-preview)
1,048,576tokens
65,536tokens
●Yes
◐Partial
◐Partial
●Yes
·
·
$2/M tok
$12/M tok
$0.2/M tok
50%
●Yes
●Yes
●Yes
◐Partial
◐Partial
Explicit caches, default 1 h TTL; implicit caching too
Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not.
–
●Yes
◐Partial
◐Partial
●Yes
2023
Long context and native multimodality across Google's cloud
Amazon BedrockMulti-vendor model access inside your existing AWS account
75.6
Multi-vendor, Claude Opus 4.8, Llama 4, Mistral, Nova Premier
1,000,000tokens
128,000tokens
●Yes
◐Partial
◐Partial
○No
·
·
–
·
·
50%
●Yes
●Yes
●Yes
●Yes
○No
5 min default, 1 h option; explicit breakpoints only
The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do.
–
●Yes
●Yes
○No
●Yes
2023
Frontier models inside your existing AWS security perimeter
Mistral AIEuropean lab with an open-weight lineage and EU-resident hosting
75.2
Mistral Large 3 (mistral-large-3-25-12)
256,000tokens
–
●Yes
◐Partial
●Yes
●Yes
·
·
$0.5/M tok
$1.5/M tok
$0.05/M tok
50%
●Yes
●Yes
◐Partial
○No
◐Partial
·
Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away.
–
●Yes
◐Partial
○No
○No
2023
European models with open weights and deployable anywhere
OpenAIGPT models plus audio, images and embeddings on one bill
73.2
GPT-5.6 Sol (gpt-5.6-sol)
1,050,000tokens
128,000tokens
●Yes
◐Partial
◐Partial
●Yes
·
·
$5/M tok
$30/M tok
$0.5/M tok
50%
●Yes
●Yes
●Yes
◐Partial
●Yes
Automatic, ≥30 min, no configuration
Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's, closer to a quarter than half a year for API snapshots, and preview models have been pulled faster still.
90days
●Yes
◐Partial
○No
-Unknown
2020
One API for text, reasoning, audio, images and embeddings
QwenAlibaba's model family, huge open-weight range, closed flagship
68.6
Qwen3.7-Max (qwen3.7-max)
1,000,000tokens
·
◐Partial
◐Partial
◐Partial
●Yes
·
·
$2.5/M tok
$7.5/M tok
$0.25/M tok
50%
●Yes
◐Partial
◐Partial
○No
◐Partial
Explicit cache 5 min (extends on hit); implicit cache no fixed TTL
Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely.
·
●Yes
-Unknown
◐Partial
◐Partial
2023
The widest open-weight family, plus a closed flagship tier
xAIGrok models with an OpenAI-shaped API and live X data access
64.5
Grok 4.5 (grok-4.5)
500,000tokens
–
●Yes
◐Partial
◐Partial
●Yes
·
·
$2/M tok
$6/M tok
$0.3/M tok
–
●Yes
●Yes
◐Partial
○No
○No
·
Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here.
–
●Yes
-Unknown
◐Partial
○No
2024
Fast, cheap frontier models with live access to X
Meta LlamaOpen-weight Llama models, hosted almost everywhere but Meta
64.3
Llama 4 Maverick
1,000,000tokens
–
●Yes
○No
●Yes
●Yes
·
·
–
·
·
·
◐Partial
◐Partial
○No
○No
○No
Host-dependent
Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you, if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears.
–
●Yes
◐Partial
○No
●Yes
2023
Open-weight models you can host anywhere, forever
Self-hosted (vLLM)baseline · Run open weights on your own GPUs behind an OpenAI-shaped API
63.8
Whatever you deploy, DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4
–
–
◐Partial
○No
●Yes
●Yes
–
·
–
·
–
0%
◐Partial
●Yes
○No
○No
○No
Automatic prefix cache in VRAM, no TTL
Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster.
–
●Yes
●Yes
○No
●Yes
2023
Your weights, your GPUs, your OpenAI-compatible endpoint
CohereEnterprise-focused models built for RAG and private deployment
63.6
Command A+ (command-a-plus-05-2026)
128,000tokens
64,000tokens
●Yes
○No
●Yes
●Yes
·
·
–
–
·
·
●Yes
●Yes
○No
○No
○No
·
Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does.
·
●Yes
◐Partial
○No
◐Partial
2021
Enterprise RAG models you can deploy inside your own network
Together AIServerless and dedicated hosting for open-weight models
62.9
Open-weight catalogue, DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4
–
–
◐Partial
◐Partial
●Yes
●Yes
·
·
–
·
·
50%
◐Partial
●Yes
○No
○No
○No
·
Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint.
–
●Yes
-Unknown
○No
○No
2023
Open-weight models with a Western contract and dedicated capacity
Fireworks AIFast open-weight inference with strong structured-output support
61.9
Open-weight catalogue, DeepSeek, Kimi K2, Qwen3, Llama 4
–
–
◐Partial
◐Partial
●Yes
●Yes
·
·
–
·
·
·
◐Partial
●Yes
○No
○No
○No
·
Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient.
·
●Yes
-Unknown
○No
○No
2023
Low-latency open-weight inference with strict structured output
AnthropicClaude models, built around long agentic runs and tool use
61.0
Claude Fable 5 (claude-fable-5)
1,000,000tokens
128,000tokens
●Yes
○No
○No
◐Partial
·
·
$10/M tok
$50/M tok
$1/M tok
50%
●Yes
●Yes
●Yes
●Yes
●Yes
5 min default, 1 h option; automatic or explicit breakpoints
Published deprecation page lists a retirement date per model, typically months ahead, Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee.
180days
◐Partial
◐Partial
○No
◐Partial
2023
Frontier models for long-horizon agentic work and code
Moonshot AIKimi models, open-weight agentic performance at low cost
60.0
Kimi K3 (kimi-k3)
1,048,576tokens
·
●Yes
○No
●Yes
●Yes
·
–
$3/M tok
$15/M tok
$0.3/M tok
–
●Yes
◐Partial
◐Partial
○No
○No
Automatic, no fixed TTL
Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence, download the checkpoint and continuity is your problem, not theirs.
·
●Yes
○No
◐Partial
○No
2023
Open-weight agentic performance at open-weight prices
OpenRouterOne OpenAI-shaped key in front of hundreds of models
59.6
Whichever upstream model you route to (400+ available)
–
–
◐Partial
◐Partial
◐Partial
●Yes
–
·
–
·
·
0%
●Yes
◐Partial
◐Partial
○No
○No
Passed through where the upstream supports caching
Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself, a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement.
–
◐Partial
◐Partial
○No
○No
2023
One key, one API shape, every model worth calling
Z.ai (GLM)GLM models with an unusually cheap flat-rate coding plan
58.2
GLM-4.6
200,000tokens
·
◐Partial
○No
●Yes
●Yes
·
·
$0.6/M tok
$2.2/M tok
·
·
●Yes
◐Partial
◐Partial
○No
○No
·
Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model.
·
●Yes
○No
◐Partial
○No
2023
Coding-focused GLM models with a flat-rate subscription
GroqCustom LPU silicon serving open-weight models at extreme speed
The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun.
–
◐Partial
-Unknown
○No
○No
2024
Custom LPU silicon for open-weight models at very low latency
CerebrasWafer-scale inference, the fastest tokens per second available
Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite.
·
◐Partial
-Unknown
○No
○No
2024
Wafer-scale inference, the highest tokens per second available
DeepSeekFrontier-adjacent models at a small fraction of Western prices
53.3
DeepSeek-V4-Flash (deepseek-v4-flash)
1,000,000tokens
384,000tokens
○No
○No
●Yes
●Yes
·
·
$0.14/M tok
$0.28/M tok
$0.0028/M tok
0%
●Yes
◐Partial
◐Partial
○No
○No
Automatic disk cache, no fixed TTL
The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases, deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently.
–
○No
○No
●Yes
○No
2023
Near-frontier quality at a fraction of the token price
Google GeminiGemini via AI Studio for prototyping or Vertex AI for production76.9
Capability
Flagship model
Gemini 3.1 Pro Preview (gemini-3.1-pro-preview)
Context window (tokens)
1,048,576tokens
Max output (tokens)
65,536tokens
Image input
●Yes
Audio in/out
◐Partial
Open weights
◐Partial
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$2/M tok
$/M output (/M tok)
$12/M tok
$/M cache read (/M tok)
$0.2/M tok
Batch discount (%)
50%
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
●Yes
Computer use
◐Partial
MCP support
◐Partial
Cache TTL
Explicit caches, default 1 h TTL; implicit caching too
Governance & continuity
Continuity policy
Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
◐Partial
Trains on your data
◐Partial
AU region
●Yes
Traction
API since
2023
Positioning
Positioning
Long context and native multimodality across Google's cloud
Amazon BedrockMulti-vendor model access inside your existing AWS account75.6
Capability
Flagship model
Multi-vendor, Claude Opus 4.8, Llama 4, Mistral, Nova Premier
Context window (tokens)
1,000,000tokens
Max output (tokens)
128,000tokens
Image input
●Yes
Audio in/out
◐Partial
Open weights
◐Partial
OpenAI-compat API
○No
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
50%
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
●Yes
Computer use
●Yes
MCP support
○No
Cache TTL
5 min default, 1 h option; explicit breakpoints only
Governance & continuity
Continuity policy
The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
●Yes
Trains on your data
○No
AU region
●Yes
Traction
API since
2023
Positioning
Positioning
Frontier models inside your existing AWS security perimeter
Mistral AIEuropean lab with an open-weight lineage and EU-resident hosting75.2
Capability
Flagship model
Mistral Large 3 (mistral-large-3-25-12)
Context window (tokens)
256,000tokens
Max output (tokens)
–
Image input
●Yes
Audio in/out
◐Partial
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$0.5/M tok
$/M output (/M tok)
$1.5/M tok
$/M cache read (/M tok)
$0.05/M tok
Batch discount (%)
50%
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
◐Partial
Computer use
○No
MCP support
◐Partial
Cache TTL
·
Governance & continuity
Continuity policy
Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
◐Partial
Trains on your data
○No
AU region
○No
Traction
API since
2023
Positioning
Positioning
European models with open weights and deployable anywhere
OpenAIGPT models plus audio, images and embeddings on one bill73.2
Capability
Flagship model
GPT-5.6 Sol (gpt-5.6-sol)
Context window (tokens)
1,050,000tokens
Max output (tokens)
128,000tokens
Image input
●Yes
Audio in/out
◐Partial
Open weights
◐Partial
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$5/M tok
$/M output (/M tok)
$30/M tok
$/M cache read (/M tok)
$0.5/M tok
Batch discount (%)
50%
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
●Yes
Computer use
◐Partial
MCP support
●Yes
Cache TTL
Automatic, ≥30 min, no configuration
Governance & continuity
Continuity policy
Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's, closer to a quarter than half a year for API snapshots, and preview models have been pulled faster still.
Notice period (days)
90days
Pinnable versions
●Yes
Zero retention
◐Partial
Trains on your data
○No
AU region
-Unknown
Traction
API since
2020
Positioning
Positioning
One API for text, reasoning, audio, images and embeddings
QwenAlibaba's model family, huge open-weight range, closed flagship68.6
Capability
Flagship model
Qwen3.7-Max (qwen3.7-max)
Context window (tokens)
1,000,000tokens
Max output (tokens)
·
Image input
◐Partial
Audio in/out
◐Partial
Open weights
◐Partial
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$2.5/M tok
$/M output (/M tok)
$7.5/M tok
$/M cache read (/M tok)
$0.25/M tok
Batch discount (%)
50%
Agentic
Tool use
●Yes
Schema output
◐Partial
Effort control
◐Partial
Computer use
○No
MCP support
◐Partial
Cache TTL
Explicit cache 5 min (extends on hit); implicit cache no fixed TTL
Governance & continuity
Continuity policy
Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely.
Notice period (days)
·
Pinnable versions
●Yes
Zero retention
-Unknown
Trains on your data
◐Partial
AU region
◐Partial
Traction
API since
2023
Positioning
Positioning
The widest open-weight family, plus a closed flagship tier
xAIGrok models with an OpenAI-shaped API and live X data access64.5
Capability
Flagship model
Grok 4.5 (grok-4.5)
Context window (tokens)
500,000tokens
Max output (tokens)
–
Image input
●Yes
Audio in/out
◐Partial
Open weights
◐Partial
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$2/M tok
$/M output (/M tok)
$6/M tok
$/M cache read (/M tok)
$0.3/M tok
Batch discount (%)
–
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
◐Partial
Computer use
○No
MCP support
○No
Cache TTL
·
Governance & continuity
Continuity policy
Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
-Unknown
Trains on your data
◐Partial
AU region
○No
Traction
API since
2024
Positioning
Positioning
Fast, cheap frontier models with live access to X
Meta LlamaOpen-weight Llama models, hosted almost everywhere but Meta64.3
Capability
Flagship model
Llama 4 Maverick
Context window (tokens)
1,000,000tokens
Max output (tokens)
–
Image input
●Yes
Audio in/out
○No
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
·
Agentic
Tool use
◐Partial
Schema output
◐Partial
Effort control
○No
Computer use
○No
MCP support
○No
Cache TTL
Host-dependent
Governance & continuity
Continuity policy
Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you, if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
◐Partial
Trains on your data
○No
AU region
●Yes
Traction
API since
2023
Positioning
Positioning
Open-weight models you can host anywhere, forever
Self-hosted (vLLM)Run open weights on your own GPUs behind an OpenAI-shaped API63.8
Capability
Flagship model
Whatever you deploy, DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial
Audio in/out
○No
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
–
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
·
$/M cache read (/M tok)
–
Batch discount (%)
0%
Agentic
Tool use
◐Partial
Schema output
●Yes
Effort control
○No
Computer use
○No
MCP support
○No
Cache TTL
Automatic prefix cache in VRAM, no TTL
Governance & continuity
Continuity policy
Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
●Yes
Trains on your data
○No
AU region
●Yes
Traction
API since
2023
Positioning
Positioning
Your weights, your GPUs, your OpenAI-compatible endpoint
CohereEnterprise-focused models built for RAG and private deployment63.6
Capability
Flagship model
Command A+ (command-a-plus-05-2026)
Context window (tokens)
128,000tokens
Max output (tokens)
64,000tokens
Image input
●Yes
Audio in/out
○No
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
–
$/M cache read (/M tok)
·
Batch discount (%)
·
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
○No
Computer use
○No
MCP support
○No
Cache TTL
·
Governance & continuity
Continuity policy
Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does.
Notice period (days)
·
Pinnable versions
●Yes
Zero retention
◐Partial
Trains on your data
○No
AU region
◐Partial
Traction
API since
2021
Positioning
Positioning
Enterprise RAG models you can deploy inside your own network
Together AIServerless and dedicated hosting for open-weight models62.9
Capability
Flagship model
Open-weight catalogue, DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial
Audio in/out
◐Partial
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
50%
Agentic
Tool use
◐Partial
Schema output
●Yes
Effort control
○No
Computer use
○No
MCP support
○No
Cache TTL
·
Governance & continuity
Continuity policy
Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint.
Notice period (days)
–
Pinnable versions
●Yes
Zero retention
-Unknown
Trains on your data
○No
AU region
○No
Traction
API since
2023
Positioning
Positioning
Open-weight models with a Western contract and dedicated capacity
Fireworks AIFast open-weight inference with strong structured-output support61.9
Capability
Flagship model
Open-weight catalogue, DeepSeek, Kimi K2, Qwen3, Llama 4
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial
Audio in/out
◐Partial
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
·
Agentic
Tool use
◐Partial
Schema output
●Yes
Effort control
○No
Computer use
○No
MCP support
○No
Cache TTL
·
Governance & continuity
Continuity policy
Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient.
Notice period (days)
·
Pinnable versions
●Yes
Zero retention
-Unknown
Trains on your data
○No
AU region
○No
Traction
API since
2023
Positioning
Positioning
Low-latency open-weight inference with strict structured output
AnthropicClaude models, built around long agentic runs and tool use61.0
Capability
Flagship model
Claude Fable 5 (claude-fable-5)
Context window (tokens)
1,000,000tokens
Max output (tokens)
128,000tokens
Image input
●Yes
Audio in/out
○No
Open weights
○No
OpenAI-compat API
◐Partial
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$10/M tok
$/M output (/M tok)
$50/M tok
$/M cache read (/M tok)
$1/M tok
Batch discount (%)
50%
Agentic
Tool use
●Yes
Schema output
●Yes
Effort control
●Yes
Computer use
●Yes
MCP support
●Yes
Cache TTL
5 min default, 1 h option; automatic or explicit breakpoints
Governance & continuity
Continuity policy
Published deprecation page lists a retirement date per model, typically months ahead, Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee.
Notice period (days)
180days
Pinnable versions
◐Partial
Zero retention
◐Partial
Trains on your data
○No
AU region
◐Partial
Traction
API since
2023
Positioning
Positioning
Frontier models for long-horizon agentic work and code
Moonshot AIKimi models, open-weight agentic performance at low cost60.0
Capability
Flagship model
Kimi K3 (kimi-k3)
Context window (tokens)
1,048,576tokens
Max output (tokens)
·
Image input
●Yes
Audio in/out
○No
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
–
Pricing
$/M input (/M tok)
$3/M tok
$/M output (/M tok)
$15/M tok
$/M cache read (/M tok)
$0.3/M tok
Batch discount (%)
–
Agentic
Tool use
●Yes
Schema output
◐Partial
Effort control
◐Partial
Computer use
○No
MCP support
○No
Cache TTL
Automatic, no fixed TTL
Governance & continuity
Continuity policy
Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence, download the checkpoint and continuity is your problem, not theirs.
Notice period (days)
·
Pinnable versions
●Yes
Zero retention
○No
Trains on your data
◐Partial
AU region
○No
Traction
API since
2023
Positioning
Positioning
Open-weight agentic performance at open-weight prices
OpenRouterOne OpenAI-shaped key in front of hundreds of models59.6
Capability
Flagship model
Whichever upstream model you route to (400+ available)
Context window (tokens)
–
Max output (tokens)
–
Image input
◐Partial
Audio in/out
◐Partial
Open weights
◐Partial
OpenAI-compat API
●Yes
Performance
TTFT p50
–
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
–
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
0%
Agentic
Tool use
●Yes
Schema output
◐Partial
Effort control
◐Partial
Computer use
○No
MCP support
○No
Cache TTL
Passed through where the upstream supports caching
Governance & continuity
Continuity policy
Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself, a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement.
Notice period (days)
–
Pinnable versions
◐Partial
Zero retention
◐Partial
Trains on your data
○No
AU region
○No
Traction
API since
2023
Positioning
Positioning
One key, one API shape, every model worth calling
Z.ai (GLM)GLM models with an unusually cheap flat-rate coding plan58.2
Capability
Flagship model
GLM-4.6
Context window (tokens)
200,000tokens
Max output (tokens)
·
Image input
◐Partial
Audio in/out
○No
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$0.6/M tok
$/M output (/M tok)
$2.2/M tok
$/M cache read (/M tok)
·
Batch discount (%)
·
Agentic
Tool use
●Yes
Schema output
◐Partial
Effort control
◐Partial
Computer use
○No
MCP support
○No
Cache TTL
·
Governance & continuity
Continuity policy
Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model.
Notice period (days)
·
Pinnable versions
●Yes
Zero retention
○No
Trains on your data
◐Partial
AU region
○No
Traction
API since
2023
Positioning
Positioning
Coding-focused GLM models with a flat-rate subscription
GroqCustom LPU silicon serving open-weight models at extreme speed55.3
The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun.
Notice period (days)
–
Pinnable versions
◐Partial
Zero retention
-Unknown
Trains on your data
○No
AU region
○No
Traction
API since
2024
Positioning
Positioning
Custom LPU silicon for open-weight models at very low latency
CerebrasWafer-scale inference, the fastest tokens per second available55.2
Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite.
Notice period (days)
·
Pinnable versions
◐Partial
Zero retention
-Unknown
Trains on your data
○No
AU region
○No
Traction
API since
2024
Positioning
Positioning
Wafer-scale inference, the highest tokens per second available
DeepSeekFrontier-adjacent models at a small fraction of Western prices53.3
Capability
Flagship model
DeepSeek-V4-Flash (deepseek-v4-flash)
Context window (tokens)
1,000,000tokens
Max output (tokens)
384,000tokens
Image input
○No
Audio in/out
○No
Open weights
●Yes
OpenAI-compat API
●Yes
Performance
TTFT p50
·
Output tok/s (tok/s)
·
Pricing
$/M input (/M tok)
$0.14/M tok
$/M output (/M tok)
$0.28/M tok
$/M cache read (/M tok)
$0.0028/M tok
Batch discount (%)
0%
Agentic
Tool use
●Yes
Schema output
◐Partial
Effort control
◐Partial
Computer use
○No
MCP support
○No
Cache TTL
Automatic disk cache, no fixed TTL
Governance & continuity
Continuity policy
The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases, deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently.
Notice period (days)
–
Pinnable versions
○No
Zero retention
○No
Trains on your data
●Yes
AU region
○No
Traction
API since
2023
Positioning
Positioning
Near-frontier quality at a fraction of the token price
●yes◐partial○no-unknownSources shown beside each value · Learn how sourcing works
measured vendor-claimed community inferredExpand any row for the source, verification date, and caveat behind every cell.
Cite this comparisonCC-BY-4.0 · verified 2026-07-30
Anthropic, OpenAI and Google lead on capability, cheaper Chinese labs win on price, routers trade features for reach; pick on continuity, not headline price. — toolweight, https://toolweight.com/compare/llm-apis, verified 2026-07-30. Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.
Frequently asked questions
What actually happens when a model I depend on is deprecated?
The alias stops resolving and requests 404
A pinned dated snapshot keeps serving until its published retirement date
A bare alias means you migrate today
Anthropic and OpenAI publish retirement dates months ahead; several providers publish nothing
Rolling aliases can change behaviour with no version bump
Is DeepSeek or Kimi really good enough to replace Claude or GPT?
For extraction, classification, summarisation, translation and bulk rewriting: yes
Roughly 1/35th the output cost of Anthropic's volume model, or 1/120th of its flagship
The honest bulk-work comparison is against the volume tier, not the flagship
For long agentic runs, sustained instruction following and first-try code, frontier labs stay ahead
Split the workload rather than picking one provider for everything
Does an OpenAI-compatible endpoint actually make switching easy?
It makes the transport identical and the semantics different
Chat completions port cleanly
Tool-call formats, effort parameters, cache accounting, thinking blocks and schema enforcement do not
Prompts tuned on one model regress on another
Budget a re-evaluation, not a config change
How much does prompt caching really save on an agent loop?
More than any other lever
Cache reads typically bill at 10-25% of the input rate
An agent loop resends the whole conversation each turn, so a long run is 80-90% cache reads
The trap: caching is a prefix match, so a timestamp or reordered tool list invalidates everything after it
Which providers can serve inference from an Australian region?
Amazon Bedrock (ap-southeast-2) and Google Vertex AI (australia-southeast1) serve frontier models in-country
First-party Anthropic, OpenAI and xAI endpoints route to US or EU by default
Residency controls vary by plan
Self-hosting is the only unambiguously AU-resident answer
Should I go through a router like OpenRouter instead of direct?
Use a router when model choice is a runtime decision, for outage fallback, or while benchmarking
Go direct when you depend on provider-specific features
That means cache breakpoints, computer use, batch endpoints, enterprise retention terms
Routers expose the intersection of upstream features, not the union
Why is the time-to-first-token column mostly empty?
Almost nobody publishes it
Self-reported latency is meaningless without a stated prompt length, region and concurrency
Figures shown are community measurements for speed-focused providers
toolweight will fill it only after running its own harness with published methodology