Frontier LLM API providers

Anthropic, OpenAI and Google lead on capability, cheaper Chinese labs win on price, routers trade features for reach; pick on continuity, not headline price.

Providers
18
Fields compared
27
Source confidence
34%
Last verified
2026-07-30 (1mo ago)
Re-verified
every 7 days
Data confidence0% of 369 figures verified in the last 7d
142sourced of 369330stale, oldest 8mo ago
18 tools · verified 1mo ago
Google Gemini logoGoogle GeminiGemini via AI Studio for prototyping or Vertex AI for production76.9

Capability

Flagship model
Gemini 3.1 Pro Preview (gemini-3.1-pro-preview)
Context window (tokens)
1,048,576tokens
Max output (tokens)
65,536tokens
Image input
Yes
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$2/M tok
$/M output (/M tok)
$12/M tok
$/M cache read (/M tok)
$0.2/M tok
Batch discount (%)
50%

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Partial
MCP support
Partial
Cache TTL
Explicit caches, default 1 h TTL; implicit caching too

Governance & continuity

Continuity policy
Stable versions carry dated suffixes you can pin, and Google publishes a retirement date per version. In practice this is the fastest-moving line-up here: preview models are withdrawn with weeks of notice, and the 1.5 generation was retired roughly a year after GA. Vertex adds contractual predictability that AI Studio does not.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
Partial
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Long context and native multimodality across Google's cloud
Amazon Bedrock logoAmazon BedrockMulti-vendor model access inside your existing AWS account75.6

Capability

Flagship model
Multi-vendor, Claude Opus 4.8, Llama 4, Mistral, Nova Premier
Context window (tokens)
1,000,000tokens
Max output (tokens)
128,000tokens
Image input
Yes
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
No

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
50%

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Yes
MCP support
No
Cache TTL
5 min default, 1 h option; explicit breakpoints only

Governance & continuity

Continuity policy
The most contractual answer in the category. Bedrock model IDs are versioned and move through a published Active to Legacy to end-of-life lifecycle with dated notices in the console and documentation, and legacy versions keep serving existing workloads after a successor lands. The cost of that stability is lag: new features reach Bedrock weeks to months after the first-party API, and some never do.
Notice period (days)
Pinnable versions
Yes
Zero retention
Yes
Trains on your data
No
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Frontier models inside your existing AWS security perimeter
Mistral AI logoMistral AIEuropean lab with an open-weight lineage and EU-resident hosting75.2

Capability

Flagship model
Mistral Large 3 (mistral-large-3-25-12)
Context window (tokens)
256,000tokens
Max output (tokens)
Image input
Yes
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$0.5/M tok
$/M output (/M tok)
$1.5/M tok
$/M cache read (/M tok)
$0.05/M tok
Batch discount (%)
50%

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Partial
Computer use
No
MCP support
Partial
Cache TTL
·

Governance & continuity

Continuity policy
Dated model IDs (the -2411 style suffix) are stable and documented, and Mistral maintains a legacy-model page listing deprecation and retirement dates. The stronger guarantee is structural: for the Apache-licensed models you can keep serving the same weights yourself after the endpoint goes away.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
European models with open weights and deployable anywhere
OpenAI logoOpenAIGPT models plus audio, images and embeddings on one bill73.2

Capability

Flagship model
GPT-5.6 Sol (gpt-5.6-sol)
Context window (tokens)
1,050,000tokens
Max output (tokens)
128,000tokens
Image input
Yes
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$5/M tok
$/M output (/M tok)
$30/M tok
$/M cache read (/M tok)
$0.5/M tok
Batch discount (%)
50%

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Partial
MCP support
Yes
Cache TTL
Automatic, ≥30 min, no configuration

Governance & continuity

Continuity policy
Dated snapshots are the norm and remain callable after the alias advances, which is the strongest routine protection in the category. Retirements are announced on a public deprecations page, but the notice window has historically been shorter than Anthropic's, closer to a quarter than half a year for API snapshots, and preview models have been pulled faster still.
Notice period (days)
90days
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
Unknown

Traction

API since
2020

Positioning

Positioning
One API for text, reasoning, audio, images and embeddings
Qwen logoQwenAlibaba's model family, huge open-weight range, closed flagship68.6

Capability

Flagship model
Qwen3.7-Max (qwen3.7-max)
Context window (tokens)
1,000,000tokens
Max output (tokens)
·
Image input
Partial
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$2.5/M tok
$/M output (/M tok)
$7.5/M tok
$/M cache read (/M tok)
$0.25/M tok
Batch discount (%)
50%

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
Partial
Cache TTL
Explicit cache 5 min (extends on hit); implicit cache no fixed TTL

Governance & continuity

Continuity policy
Model Studio exposes dated snapshot IDs alongside rolling aliases, so pinning is possible if you use the dated form. Deprecation notices appear in the console rather than as a public policy page, and the cadence of new Qwen releases means aliases move quickly. Open-weight tiers remove the problem entirely.
Notice period (days)
·
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
Partial
AU region
Partial

Traction

API since
2023

Positioning

Positioning
The widest open-weight family, plus a closed flagship tier
xAI logoxAIGrok models with an OpenAI-shaped API and live X data access64.5

Capability

Flagship model
Grok 4.5 (grok-4.5)
Context window (tokens)
500,000tokens
Max output (tokens)
Image input
Yes
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$2/M tok
$/M output (/M tok)
$6/M tok
$/M cache read (/M tok)
$0.3/M tok
Batch discount (%)

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
Dated model IDs exist and are the recommended way to address a model, but xAI publishes no formal deprecation policy or notice window, and model naming has changed repeatedly between generations. The shortest track record of any first-party lab here.
Notice period (days)
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
Partial
AU region
No

Traction

API since
2024

Positioning

Positioning
Fast, cheap frontier models with live access to X
Meta Llama logoMeta LlamaOpen-weight Llama models, hosted almost everywhere but Meta64.3

Capability

Flagship model
Llama 4 Maverick
Context window (tokens)
1,000,000tokens
Max output (tokens)
Image input
Yes
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
·

Agentic

Tool use
Partial
Schema output
Partial
Effort control
No
Computer use
No
MCP support
No
Cache TTL
Host-dependent

Governance & continuity

Continuity policy
Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you, if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears.
Notice period (days)
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Open-weight models you can host anywhere, forever
Self-hosted (vLLM) logoSelf-hosted (vLLM)Run open weights on your own GPUs behind an OpenAI-shaped API63.8

Capability

Flagship model
Whatever you deploy, DeepSeek-V3.x, Qwen3, Kimi K2, Llama 4
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
Batch discount (%)
0%

Agentic

Tool use
Partial
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL
Automatic prefix cache in VRAM, no TTL

Governance & continuity

Continuity policy
Permanent, and this is the whole point of the baseline. The checkpoint sits on your disk. Nobody can deprecate it, re-price it, silently upgrade it, or change its behaviour between releases. The obligation transfers to you: you own the GPUs, the upgrades, the security patches and the on-call roster.
Notice period (days)
Pinnable versions
Yes
Zero retention
Yes
Trains on your data
No
AU region
Yes

Traction

API since
2023

Positioning

Positioning
Your weights, your GPUs, your OpenAI-compatible endpoint
Cohere logoCohereEnterprise-focused models built for RAG and private deployment63.6

Capability

Flagship model
Command A+ (command-a-plus-05-2026)
Context window (tokens)
128,000tokens
Max output (tokens)
64,000tokens
Image input
Yes
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
$/M cache read (/M tok)
·
Batch discount (%)
·

Agentic

Tool use
Yes
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does.
Notice period (days)
·
Pinnable versions
Yes
Zero retention
Partial
Trains on your data
No
AU region
Partial

Traction

API since
2021

Positioning

Positioning
Enterprise RAG models you can deploy inside your own network
Together AI logoTogether AIServerless and dedicated hosting for open-weight models62.9

Capability

Flagship model
Open-weight catalogue, DeepSeek-V3.x, Kimi K2, Qwen3, Llama 4
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
50%

Agentic

Tool use
Partial
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
Serverless models are rotated off when demand drops, typically with a deprecations page and short notice measured in weeks. Because everything served is open-weight, the mitigation is always available: move to a dedicated endpoint on the same platform, to another host, or onto your own GPUs with the same checkpoint.
Notice period (days)
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
Open-weight models with a Western contract and dedicated capacity
Fireworks AI logoFireworks AIFast open-weight inference with strong structured-output support61.9

Capability

Flagship model
Open-weight catalogue, DeepSeek, Kimi K2, Qwen3, Llama 4
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
·

Agentic

Tool use
Partial
Schema output
Yes
Effort control
No
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
Same shape as Together: a published deprecations list, serverless models retired on weeks of notice when demand falls, and dedicated deployments as the stable option. Open weights throughout mean no retirement is terminal, only inconvenient.
Notice period (days)
·
Pinnable versions
Yes
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
Low-latency open-weight inference with strict structured output
Anthropic logoAnthropicClaude models, built around long agentic runs and tool use61.0

Capability

Flagship model
Claude Fable 5 (claude-fable-5)
Context window (tokens)
1,000,000tokens
Max output (tokens)
128,000tokens
Image input
Yes
Audio in/out
No
Open weights
No
OpenAI-compat API
Partial

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$10/M tok
$/M output (/M tok)
$50/M tok
$/M cache read (/M tok)
$1/M tok
Batch discount (%)
50%

Agentic

Tool use
Yes
Schema output
Yes
Effort control
Yes
Computer use
Yes
MCP support
Yes
Cache TTL
5 min default, 1 h option; automatic or explicit breakpoints

Governance & continuity

Continuity policy
Published deprecation page lists a retirement date per model, typically months ahead, Opus 3 went on 2026-01-05, Sonnet 3.7 and Haiku 3.5 on 2026-02-19. Older models carried dated snapshot IDs you could pin, but the current generation ships as bare aliases with no dated ID behind them, so the retirement date is your only guarantee.
Notice period (days)
180days
Pinnable versions
Partial
Zero retention
Partial
Trains on your data
No
AU region
Partial

Traction

API since
2023

Positioning

Positioning
Frontier models for long-horizon agentic work and code
Moonshot AI logoMoonshot AIKimi models, open-weight agentic performance at low cost60.0

Capability

Flagship model
Kimi K3 (kimi-k3)
Context window (tokens)
1,048,576tokens
Max output (tokens)
·
Image input
Yes
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)

Pricing

$/M input (/M tok)
$3/M tok
$/M output (/M tok)
$15/M tok
$/M cache read (/M tok)
$0.3/M tok
Batch discount (%)

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
Automatic, no fixed TTL

Governance & continuity

Continuity policy
Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence, download the checkpoint and continuity is your problem, not theirs.
Notice period (days)
·
Pinnable versions
Yes
Zero retention
No
Trains on your data
Partial
AU region
No

Traction

API since
2023

Positioning

Positioning
Open-weight agentic performance at open-weight prices
OpenRouter logoOpenRouterOne OpenAI-shaped key in front of hundreds of models59.6

Capability

Flagship model
Whichever upstream model you route to (400+ available)
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Partial
OpenAI-compat API
Yes

Performance

TTFT p50
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
0%

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
Passed through where the upstream supports caching

Governance & continuity

Continuity policy
Structurally the best hedge on this page and formally the weakest promise. OpenRouter guarantees nothing itself, a closed model vanishes the moment its lab retires it. But for open-weight models the same ID stays routable because several independent hosts serve it, and automatic failover means a single host going dark is invisible to your application. Buy it as insurance against host failure, not against model retirement.
Notice period (days)
Pinnable versions
Partial
Zero retention
Partial
Trains on your data
No
AU region
No

Traction

API since
2023

Positioning

Positioning
One key, one API shape, every model worth calling
Z.ai (GLM) logoZ.ai (GLM)GLM models with an unusually cheap flat-rate coding plan58.2

Capability

Flagship model
GLM-4.6
Context window (tokens)
200,000tokens
Max output (tokens)
·
Image input
Partial
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$0.6/M tok
$/M output (/M tok)
$2.2/M tok
$/M cache read (/M tok)
·
Batch discount (%)
·

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
Versioned model names (GLM-4.5, 4.6) stay addressable after a successor ships, but there is no published deprecation policy or notice period. MIT-licensed weights are the fallback: every flagship generation so far has been released publicly, so a retired endpoint does not strand the model.
Notice period (days)
·
Pinnable versions
Yes
Zero retention
No
Trains on your data
Partial
AU region
No

Traction

API since
2023

Positioning

Positioning
Coding-focused GLM models with a flat-rate subscription
Groq logoGroqCustom LPU silicon serving open-weight models at extreme speed55.3

Capability

Flagship model
Curated open-weight models (Kimi K2, Llama 4, Qwen3, GPT-OSS)
Context window (tokens)
Max output (tokens)
Image input
Partial
Audio in/out
Partial
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
250ms
Output tok/s (tok/s)
400tok/s

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)

Agentic

Tool use
Partial
Schema output
Partial
Effort control
No
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
The weakest continuity of any host here. Models are added and removed from the console on a short cycle as Groq re-allocates LPU capacity, deprecations are announced with weeks rather than months of notice, and there is no dedicated-endpoint escape hatch for a model that gets pulled. Design for substitution: keep two model IDs configured and an eval you can rerun.
Notice period (days)
Pinnable versions
Partial
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2024

Positioning

Positioning
Custom LPU silicon for open-weight models at very low latency
Cerebras logoCerebrasWafer-scale inference, the fastest tokens per second available55.2

Capability

Flagship model
Curated open-weight models (Qwen3, GLM, Llama, GPT-OSS)
Context window (tokens)
Max output (tokens)
Image input
No
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
200ms
Output tok/s (tok/s)
2,000tok/s

Pricing

$/M input (/M tok)
$/M output (/M tok)
·
$/M cache read (/M tok)
·
Batch discount (%)
·

Agentic

Tool use
Partial
Schema output
Partial
Effort control
No
Computer use
No
MCP support
No
Cache TTL
·

Governance & continuity

Continuity policy
Same fragility as Groq, from the same cause: a small curated menu sized to available wafer-scale capacity, with models added and removed as demand shifts. No published notice window. The models themselves are open-weight, so a removal costs you a re-benchmark and a host change rather than a rewrite.
Notice period (days)
·
Pinnable versions
Partial
Zero retention
Unknown
Trains on your data
No
AU region
No

Traction

API since
2024

Positioning

Positioning
Wafer-scale inference, the highest tokens per second available
DeepSeek logoDeepSeekFrontier-adjacent models at a small fraction of Western prices53.3

Capability

Flagship model
DeepSeek-V4-Flash (deepseek-v4-flash)
Context window (tokens)
1,000,000tokens
Max output (tokens)
384,000tokens
Image input
No
Audio in/out
No
Open weights
Yes
OpenAI-compat API
Yes

Performance

TTFT p50
·
Output tok/s (tok/s)
·

Pricing

$/M input (/M tok)
$0.14/M tok
$/M output (/M tok)
$0.28/M tok
$/M cache read (/M tok)
$0.0028/M tok
Batch discount (%)
0%

Agentic

Tool use
Yes
Schema output
Partial
Effort control
Partial
Computer use
No
MCP support
No
Cache TTL
Automatic disk cache, no fixed TTL

Governance & continuity

Continuity policy
The worst on this page for a closed endpoint, and the best if you self-host. The API offers only rolling aliases, deepseek-chat silently points at whatever the current model is, so behaviour can change with no version to pin and no deprecation notice. The mitigation is the MIT licence: download the checkpoint and the version is yours permanently.
Notice period (days)
Pinnable versions
No
Zero retention
No
Trains on your data
Yes
AU region
No

Traction

API since
2023

Positioning

Positioning
Near-frontier quality at a fraction of the token price
yespartialnounknown
Sources shown beside each value · Learn how sourcing works
measured vendor-claimed community inferredExpand any row for the source, verification date, and caveat behind every cell.

Where each one wins

  • Workloads that genuinely need 500K+ tokens of context per request
  • Voice and video products wanting one model rather than a pipeline
  • Australian or regulated deployments that must pin inference to a region

The cheapest million-token window and only genuinely native audio, but the fastest-churning line-up here, so expect yearly migrations.

  • Regulated workloads that must stay inside an existing AWS perimeter
  • Australian data residency without self-hosting
  • Teams that need a documented model lifecycle to satisfy an auditor

Choose Bedrock for procurement and governance, not capability; accept lagging features and no MCP connector or automatic caching.

  • EU data-residency requirements that rule out US-hosted inference
  • On-premise deployment with a commercial support contract
  • Teams that want an open-weight escape hatch under the hosted API

Rarely the capability leader, consistently the answer when jurisdiction is the blocker; check licences per model.

OpenAI

73.2
  • Multimodal products that need speech and image generation alongside text
  • Teams standardising on one vendor for everything rather than best-of-breed
  • Anyone who wants a dated snapshot they can actually pin

The widest surface area, with audio, images and text on one bill; pinnable snapshots but shorter observed notice.

Qwen

68.6
  • Finding an open-weight model at a specific size or modality
  • Fine-tuning where a permissive licence and a model-size ladder both matter
  • Asia-Pacific deployments already on Alibaba Cloud

The best open-weight range, a model at nearly every size and modality; the hosted Max tier is rarely the reason to be here.

xAI

64.5
  • Products that want real-time X/Twitter context in the model
  • Cheap evaluation of a frontier alternative with a one-line base-URL change
  • Consumer applications where governance requirements are light

Easy to try and strong on price-performance, but the thinnest governance here, so fine for consumer, hard for regulated work.

  • Workloads that must survive any single vendor disappearing
  • On-premise or air-gapped deployments with a recognisable licence
  • Fine-tuning on your own data without a vendor in the loop

A model family, not a serious first-party endpoint; its value is continuity insurance, a model that cannot be deprecated.

  • Data that legally cannot leave your infrastructure
  • Sustained high-utilisation inference where GPU-hours beat per-token billing
  • Guaranteeing a model version survives indefinitely

A reference point, not a recommendation; it wins on cost only at high steady load, so choose it for sovereignty or permanence.

Cohere

63.6
  • RAG pipelines where rerank quality drives the result
  • Deployments that must run inside your own VPC or data centre
  • Enterprises whose procurement needs a vendor agreement, not a credit card

Competes on retrieval, private deployment and contracts; its rerank and embeddings are the real draw, the chat API overpriced.

  • Running DeepSeek or Kimi weights under a Western contract
  • Steady high-volume load that justifies a dedicated GPU endpoint
  • Fine-tuning open-weight models without owning hardware

The pragmatic way to run Chinese open weights on US infrastructure under a recognisable contract; dedicated endpoints shine at steady load.

  • Open-weight workloads where output must validate against a schema every time
  • Latency-sensitive serving without moving to custom silicon
  • Fine-tuned open models served on managed infrastructure

Hard to separate from Together, but stronger on constrained decoding and latency; pick it when output must parse every time.

  • Agent harnesses that call tools for minutes at a time
  • Code generation and review where first-try correctness pays for the token price
  • Teams that want cost tuned per route via graduated effort levels

The best platform for long tool loops, but continuity regressed to retirement dates and Fable 5 can refuse outright.

  • Agent loops on an open-weight budget
  • Self-hosting a tool-use-capable model under a permissive licence
  • Replacing a frontier model on the cheaper half of a mixed workload

Kimi K2 is the strongest case that open weights caught up on agentic work; reach it through Groq, Fireworks or Together.

  • Keeping model choice a runtime decision rather than an architectural one
  • Benchmarking a dozen models without a dozen accounts
  • Failover when a single upstream host goes down

The highest-leverage integration: one endpoint, hundreds of models, failover; the trade is intersection features, not the union.

  • Driving a coding agent all day on a predictable monthly cost
  • Cheap high-volume code completion and refactoring
  • Self-hosting an MIT-licensed coding model

The flat-rate coding plan is genuinely disruptive, turning all-day agent use into a fixed cost; quality trails only on hard reasoning.

Groq

55.3
  • Voice agents and any interface where perceived latency is the feature
  • High-volume cheap inference on well-known open-weight models
  • Fast transcription alongside text generation on one key

When instant responses are the product, Groq changes what you can build; not general-purpose, with a short, rotating menu.

  • Reasoning models where thinking-token latency is the bottleneck
  • Interactive tools that feel broken at 50 tokens per second
  • Benchmarking how much of your UX problem is actually latency

The fastest single-stream generation you can buy, by a qualitative margin; the catch is Groq's: short menu, no continuity guarantees.

  • High-volume classification, extraction and summarisation
  • Cost-sensitive workloads where a 35x output-price gap against a volume-tier frontier model outweighs a small quality gap
  • Self-hosting a capable model under a genuinely permissive licence

The category's price floor with unmatched quality-per-dollar for bulk work; for regulated data, use the MIT weights via a Western host.

Which frontier LLM API should I build on?

  • Every provider quotes $/M tokens; almost none tell you how long the model lives
  • That continuity gap is the biggest hidden cost, so this page leads with it
  • AnthropicAnthropic logo for agentic work: effort control, prompt caching and MCP are first-class
  • OpenAIOpenAI logo for audio, image generation and text on one bill, or Responses API fluency
  • Google for cheap million-token context, native audio, or AU-resident Vertex deployment
  • Those three are not interchangeable; a tuned prompt lands differently on each
  • Chinese open-weight labs (DeepSeekDeepSeek logo, Moonshot, QwenQwen logo, Z.ai) reset the price floor
  • DeepSeekDeepSeek logo bills $0.42/M output vs $50 for Claude Fable 5, ~120x cheaper
  • Still ~35x cheaper than the volume tier most teams would actually use
  • For classification, extraction, summarisation and bulk rewriting, the quality gap is not worth the price gap
  • The catch is governance: permissive retention, no ZDR, rolling aliases that upgrade under you
  • For sensitive data, self-host the weights or route through a contractually clean host
  • OpenRouterOpenRouter logo is the cheapest insurance: one integration, hundreds of models, retired models often still reachable
  • GroqGroq logo and CerebrasCerebras logo are latency instruments, not general-purpose substitutes
  • Worth an entire architecture for voice, autocomplete or interactive search; frustrating otherwise
  • Prices keep falling; context windows have stopped being the differentiator
  • The fight has moved to agentic fidelity: parallel tool calls, strict schema, durable caching
  • Assume today's model retires within two years
  • Build the abstraction layer now, keep a re-runnable eval set, treat every provider as replaceable

The pricing traps nobody puts on the pricing page

  • Headline $/M is the least useful number here; four things distort it
  • Output tokens dominate: a reasoning model burns tens of thousands of thinking tokens, billed at the output rate
  • A cheap-input, expensive-output provider is dear for reasoning, cheap for RAG answering
  • Model your actual input:output ratio before comparing
  • Cached input is where the money is: cache reads bill at ~10% of input at AnthropicAnthropic logo and OpenAIOpenAI logo
  • An agent loop resending its history is mostly cache reads
  • Caching is a prefix match; a timestamp, non-deterministic JSON, or a per-user tool list destroys it silently
  • Watch the cache-read token counter, not the invoice
  • Long-context surcharges: Google charges more above 200K tokens; others meter cache storage per hour
  • A million-token window at the short-context rate is not what you pay
  • Batch, but check the column first: where an async endpoint exists the discount is usually a flat 50%
  • AnthropicAnthropic logo, OpenAIOpenAI logo, Google, Mistral, QwenQwen logo, Bedrock and Together all publish that rate
  • DeepSeekDeepSeek logo offers no batch endpoint; its list price already sits below most batch rates
  • OpenRouterOpenRouter logo offers none either, a real cost if half your workload tolerates async
  • xAIxAI logo, Moonshot, Z.ai, CohereCohere logo, Fireworks and CerebrasCerebras logo publish nothing confirmable
  • GroqGroq logo documents a batch discount without a percentage we could stand behind
  • Where the halving exists and latency is negotiable, no negotiation beats it; where it is not, no committed spend conjures it

How to make a migration survivable

  • Assume an eighteen-month life for whatever you ship on
  • Providers that publish retirement dates help you; those that don't will break you
  • Pin a dated snapshot wherever offered, and record the pin in configuration, not code
  • This is getting harder: AnthropicAnthropic logo's current line-up ships bare aliases, so you fall back on the retirement date
  • Keep the provider behind one seam: a per-provider module owning request construction, retries and parsing
  • Not an abstraction that pretends every model is the same; that fails on thinking blocks or tool-result shapes
  • Two implementations behind one interface is cheap insurance; a universal adapter is expensive fantasy
  • Keep a re-runnable eval set: 50-200 real inputs with graded expected outputs
  • Without it a forced migration is a multi-week vibe check; with it, a morning's work and a defensible number
  • Know your escape hatch first: for open weights it is total, vLLM serves them as long as you have GPUs
  • For closed models it is a router that may still have an upstream, or nothing at all

Deciding: three workloads, three answers

  • Bulk text processing (classification, extraction, tagging, translation, firehose summarising): price dominates, capability barely registers
  • DeepSeekDeepSeek logo, QwenQwen logo, Moonshot and Z.ai are the rational choices
  • Or GroqGroq logo and Together for open weights with a Western contract
  • Add the batch endpoint if latency is negotiable
  • Never send sensitive data through a provider whose retention terms you haven't read
  • Agent harnesses (plan, call tools, read results, iterate for minutes): this is where frontier labs earn their price
  • You need well-formed tool calls at depth, enforced schema output, effort control, and caching that survives the loop
  • AnthropicAnthropic logo and OpenAIOpenAI logo are credible; Google is close and cheaper on long context
  • Budget for a long agentic turn running for minutes on one request
  • Interactive, latency-critical UI (voice, live search, inline completion): time to first token is the product
  • CerebrasCerebras logo and GroqGroq logo are an order of magnitude ahead of first-party endpoints on tok/s
  • The constraint: you take whichever open-weight models they host
  • Design the feature around the model menu, not the reverse

What is frontier llm apis compared, price, context, tool use and deprecation policy?

Frontier LLM APIs compared, price, context, tool use and deprecation policy

On toolweight, Frontier LLM APIs compared, price, context, tool use and deprecation policy means the 18 tools benchmarked on this page, AnthropicAnthropic logo, OpenAIOpenAI logo, Google GeminiGoogle Gemini logo, xAIxAI logo, Meta LlamaMeta Llama logo, Mistral AIMistral AI logo, DeepSeekDeepSeek logo, QwenQwen logo, Moonshot AIMoonshot AI logo, Z.ai (GLM)Z.ai (GLM) logo, Amazon BedrockAmazon Bedrock logo, CohereCohere logo, OpenRouterOpenRouter logo, Together AITogether AI logo, Fireworks AIFireworks AI logo, GroqGroq logo, CerebrasCerebras logo, Self-hosted (vLLM)Self-hosted (vLLM) logo, judged on the same 27 fields, from the same sources, on the same date. The question it exists to answer: Which frontier LLM API should I build on?

How does toolweight compare these?

  • Pricing here is stale within a fortnight; this category re-prices faster than any other on toolweight
  • Seven-day verification cadence, every price cell dated
  • Treat any cell dated more than a month ago as indicative; click through to the vendor first
  • Figures are pay-as-you-go list rates in USD per million tokens
  • Before batch, cache or committed-spend discounts, and before long-context surcharges
  • Where a vendor runs an unexpired introductory rate, the column carries list price and the note carries the intro rate and expiry
  • An introductory rate is what you actually pay today, so read the note before modelling a bill
  • Entries are providers, not models
  • Every price, context and capability figure describes that provider's current flagship, named in the Flagship model column
  • A cheaper flagship often sits beside a cheaper mid-tier that would serve you better
  • The killer column records notice before retirement, and whether you can pin a dated snapshot
  • Continuity is judged from published deprecation pages and observed retirements, not marketing
  • Where a provider publishes no policy, we say so rather than guess
  • Latency figures are community-reported; toolweight has not run its own TTFT harness, so most of that column is blank
Full methodology and sourcing policy →
Cite this comparisonCC-BY-4.0 · verified 2026-07-30

Anthropic, OpenAI and Google lead on capability, cheaper Chinese labs win on price, routers trade features for reach; pick on continuity, not headline price. — toolweight, https://toolweight.com/compare/llm-apis, verified 2026-07-30. Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

Frequently asked questions

What actually happens when a model I depend on is deprecated?

  • The alias stops resolving and requests 404
  • A pinned dated snapshot keeps serving until its published retirement date
  • A bare alias means you migrate today
  • AnthropicAnthropic logo and OpenAIOpenAI logo publish retirement dates months ahead; several providers publish nothing
  • Rolling aliases can change behaviour with no version bump

Is DeepSeek or Kimi really good enough to replace Claude or GPT?

  • For extraction, classification, summarisation, translation and bulk rewriting: yes
  • Roughly 1/35th the output cost of AnthropicAnthropic logo's volume model, or 1/120th of its flagship
  • The honest bulk-work comparison is against the volume tier, not the flagship
  • For long agentic runs, sustained instruction following and first-try code, frontier labs stay ahead
  • Split the workload rather than picking one provider for everything

Does an OpenAI-compatible endpoint actually make switching easy?

  • It makes the transport identical and the semantics different
  • Chat completions port cleanly
  • Tool-call formats, effort parameters, cache accounting, thinking blocks and schema enforcement do not
  • Prompts tuned on one model regress on another
  • Budget a re-evaluation, not a config change

How much does prompt caching really save on an agent loop?

  • More than any other lever
  • Cache reads typically bill at 10-25% of the input rate
  • An agent loop resends the whole conversation each turn, so a long run is 80-90% cache reads
  • The trap: caching is a prefix match, so a timestamp or reordered tool list invalidates everything after it

Which providers can serve inference from an Australian region?

  • Amazon BedrockAmazon Bedrock logo (ap-southeast-2) and Google Vertex AI (australia-southeast1) serve frontier models in-country
  • First-party AnthropicAnthropic logo, OpenAIOpenAI logo and xAIxAI logo endpoints route to US or EU by default
  • Residency controls vary by plan
  • Self-hosting is the only unambiguously AU-resident answer

Should I go through a router like OpenRouter instead of direct?

  • Use a router when model choice is a runtime decision, for outage fallback, or while benchmarking
  • Go direct when you depend on provider-specific features
  • That means cache breakpoints, computer use, batch endpoints, enterprise retention terms
  • Routers expose the intersection of upstream features, not the union

Why is the time-to-first-token column mostly empty?

  • Almost nobody publishes it
  • Self-reported latency is meaningless without a stated prompt length, region and concurrency
  • Figures shown are community measurements for speed-focused providers
  • toolweight will fill it only after running its own harness with published methodology