---
title: Cheapest per token
slug: best-llm-api-for-coding-agents
url: "https://toolweight.com/best/best-llm-api-for-coding-agents"
category: llm-apis
preset_weights: "price_in:10,price_out:10,cached_input:6,batch_discount:4,context_window:2,open_weights:2"
last_verified: 2026-06-24
license: CC-BY-4.0
---

# Cheapest per token

> Ranks on flagship list price, with cache and batch discounts weighted in. Rows that publish no flat per-token rate for a named flagship, the routers, the multi-model catalogues and the self-host baseline, are marked inapplicable rather than ranked, because 'cheapest' is not a question they answer.

**OpenAI** ranks first for this intent at 81.9 / 100, scored across 23 weighted fields from the LLM APIs registry. GPT models plus audio, images and embeddings on one bill

## Ranking

| # | Tool | Score | Coverage | Why |
| --- | --- | --- | --- | --- |
| 1 | [OpenAI](https://toolweight.com/options/openai) | 81.9 | 87% | $/M input 88%, $/M output 80%, Tool use 100% |
| 2 | [Google Gemini](https://toolweight.com/options/google-gemini) | 79.7 | 83% | $/M input 80%, Tool use 100%, $/M output 76% |
| 3 | [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock) | 77.9 | 74% | Tool use 100%, Schema output 100%, Pinnable versions 100% |
| 4 | [Mistral AI](https://toolweight.com/options/mistral-ai) | 73.7 | 65% | Tool use 100%, Schema output 100%, Pinnable versions 100% |
| 5 | [Qwen](https://toolweight.com/options/alibaba-qwen) | 69.4 | 74% | $/M input 88%, $/M output 88%, Tool use 100% |
| 6 | [Meta Llama](https://toolweight.com/options/meta-llama) | 65.5 | 65% | Pinnable versions 100%, Trains on your data 100%, OpenAI-compat API 100% |
| 7 | [Self-hosted (vLLM)](https://toolweight.com/options/self-hosted-vllm) | 65.3 | 65% | Schema output 100%, Pinnable versions 100%, Trains on your data 100% |
| 8 | [Cohere](https://toolweight.com/options/cohere) | 64.6 | 74% | $/M output 80%, Tool use 100%, $/M input 75% |
| 9 | [xAI](https://toolweight.com/options/xai) | 64.5 | 57% | Tool use 100%, Schema output 100%, Pinnable versions 100% |
| 10 | [Together AI](https://toolweight.com/options/together-ai) | 64.3 | 61% | Schema output 100%, Pinnable versions 100%, Trains on your data 100% |
| 11 | [Fireworks AI](https://toolweight.com/options/fireworks-ai) | 63.3 | 57% | Schema output 100%, Pinnable versions 100%, Trains on your data 100% |
| 12 | [OpenRouter](https://toolweight.com/options/openrouter) | 61.7 | 65% | Tool use 100%, Trains on your data 100%, OpenAI-compat API 100% |
| 13 | [Z.ai (GLM)](https://toolweight.com/options/zhipu-zai) | 61.4 | 74% | $/M output 96%, $/M input 94%, Tool use 100% |
| 14 | [Anthropic](https://toolweight.com/options/anthropic) | 61.2 | 91% | Tool use 100%, Schema output 100%, Notice period 100% |
| 15 | [Moonshot AI](https://toolweight.com/options/moonshot-ai) | 60.2 | 78% | $/M output 95%, $/M input 94%, Tool use 100% |
| 16 | [Groq](https://toolweight.com/options/groq) | 56.7 | 65% | Trains on your data 100%, OpenAI-compat API 100%, Tool use 60% |
| 17 | [Cerebras](https://toolweight.com/options/cerebras) | 56.5 | 65% | Trains on your data 100%, OpenAI-compat API 100%, Tool use 60% |
| 18 | [DeepSeek](https://toolweight.com/options/deepseek) | 47.8 | 83% | $/M output 99%, $/M input 97%, $/M cache read 97% |

## Weights used

| Field | Weight | Direction |
| --- | --- | --- |
| $/M input | 10 | lower is better |
| $/M output | 10 | lower is better |
| Tool use | 8 | higher is better |
| Schema output | 7 | higher is better |
| Notice period | 7 | higher is better |
| Pinnable versions | 7 | higher is better |
| Trains on your data | 7 | lower is better |
| $/M cache read | 6 | lower is better |
| Effort control | 6 | higher is better |
| Zero retention | 6 | higher is better |
| OpenAI-compat API | 5 | higher is better |
| MCP support | 5 | higher is better |
| Image input | 4 | higher is better |
| TTFT p50 | 4 | lower is better |
| Output tok/s | 4 | higher is better |
| Batch discount | 4 | higher is better |
| AU region | 4 | higher is better |
| Max output | 3 | higher is better |
| Audio in/out | 3 | higher is better |
| Computer use | 3 | higher is better |
| Context window | 2 | higher is better |
| Open weights | 2 | higher is better |
| API since | 1 | lower is better |

## Verdict

Every provider on this page will quote you a $/M token figure. Almost none of them will tell you how long the model behind it lives. That asymmetry is the single biggest hidden cost in this category, and it is why this page leads with a continuity column rather than a price column.

For most production work in 2026 the shortlist is short. Anthropic if the workload is agentic, tool calls, long multi-step runs, code, because effort control, prompt caching and MCP are first-class rather than bolted on. OpenAI if you need audio, image generation and text behind one billing relationship, or if your team is already fluent in the Responses API. Google if you need million-token context cheaply, native audio in and out, or an AU-resident deployment through Vertex. Those three are not interchangeable at the prompt level: a prompt tuned on one lands differently on the others, and the effort/thinking knobs have no common vocabulary.

The Chinese open-weight labs, DeepSeek, Moonshot, Qwen, Z.ai, have changed what the floor looks like. DeepSeek's flagship bills $0.42 per million output tokens against $50 for Claude Fable 5 and $15 for Claude Sonnet 5, roughly a hundred and twenty times cheaper than Anthropic's top-end model, and still around thirty-five times cheaper than the volume tier most teams would actually reach for. For classification, extraction, summarisation and bulk rewriting the quality gap does not justify the price gap. The catch is governance, not capability: retention defaults are permissive, zero-data-retention is generally not on offer, and several expose only a rolling alias that silently upgrades under you. If the data is sensitive, run the weights yourself or route through a host that is contractually clean rather than calling the origin API.

Routers deserve more credit than they get. OpenRouter is the cheapest insurance policy in the category, one integration, hundreds of models, and a real chance that a retired model stays reachable somewhere because an open-weight copy is still hosted. Groq and Cerebras are not general-purpose substitutes; they are latency instruments. If your product's differentiator is that the answer appears instantly, voice, autocomplete, interactive search, they are worth an entire architecture. If it is not, their model menu will frustrate you within a quarter.

Where this is heading: prices keep falling, context windows have stopped being the differentiator, and the fight has moved to agentic fidelity, parallel tool calls that do not degrade, structured output that never breaks schema, caching that survives a long agent loop. Assume the model you ship on today will be retired inside two years. Build the abstraction layer now, keep an evaluation set you can re-run in an afternoon, and treat every provider on this page as replaceable.

Full roster and provenance: [LLM APIs](https://toolweight.com/compare/llm-apis) · machine-readable at https://toolweight.com/compare/llm-apis.md

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/best/best-llm-api-for-coding-agents
- Machine-readable: https://toolweight.com/best/best-llm-api-for-coding-agents.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
