DeepSeek

DeepSeek is frontier-adjacent models at a small fraction of Western prices. We compare it in LLM APIs. It is top of its roster on $/M output ($0.42 /M tok). The trade-off is being marked No on Pinnable versions, where OpenAI records Yes.

What is DeepSeek?

DeepSeek — Frontier-adjacent models at a small fraction of Western prices. DeepSeek reset the category's price floor and has kept cutting since. The V3 line is released under MIT, so the same model you call over the API can be run on your own hardware or through a Western host. The first-party API is the cheapest serious option on this page; its governance terms are also the most permissive, which is the reason to route around it for sensitive workloads. Founded 2023. Open source under MIT.

We track it in 1 comparison — LLM APIs — so every claim below is a cell in a table you can open and check rather than an impression. Across those rosters it sits against 17 other tools, and what follows is where it visibly separates from them.

Where it wins.

  • $/M output — $0.42 /M tok, the best figure of the 8 here.
  • $/M input — $0.28 /M tok, the best figure of the 8 here.
  • Trains on your data — Yes, and it is the only one of 18 here that manages it.
  • $/M cache read — $0.028 /M tok, the best figure of the 5 here.

Each of those is ranked against the whole roster on its category page, not against a hand-picked subset, so a first place here means first of everything we list.

Where it gives ground.

  • Pinnable versions — No. OpenAI records Yes.
  • Zero retention — No. Amazon Bedrock records Yes.
  • Context window — 128,000 tokens, the weakest of the 10 here. Anthropic records 1,000,000 tokens.
  • Image input — No. Anthropic records Yes.

None of these disqualify it on their own. They are the fields to check against your own requirements before you commit, because they are the ones where a competitor genuinely does better.

Provenance. 23 of 27 tracked fields carry a value for DeepSeek, and 4 of those cite a document you can open. Last verified 2026-01-15. Every figure keeps its own provenance — measured by us, claimed by the vendor, inferred, or community-reported — and we would rather print a dash than a guess.

Its nearest neighbour in our data is Qwen. DeepSeek is ahead on Trains on your data (Yes against Partial) and Open weights (Yes against Partial). Qwen takes Pinnable versions (Yes against No) and Batch discount (50 % against 0 %). That pattern repeats across the rest of the roster — see DeepSeek alternatives for the other rivals, each compared the same way.

At a glance
Founded
2023
Licence
MIT (open source)
Fields we track
23 of 27
Last verified
2026-01-15

DeepSeek in LLM APIs

Ranked against 18 tools across 27 sourced fields. Open the full LLM APIs table.

The price floor of the category and the reason everyone else's rates fell. For bulk text work the quality-per-dollar is unmatched. Do not send regulated data to the first-party endpoint — permissive retention terms and rolling aliases with no pinning make it unsuitable for anything sensitive or long-lived. Use the MIT weights through a Western host instead.

Where it lands in this roster
$/M output
$0.42 /M tok1st of 8
Inferredverified 2026-01-15source
$/M input
$0.28 /M tok1st of 8
Inferredverified 2026-01-15source
Pinnable versions
No18th of 18
Community-reportedverified 2026-01-15
Trains on your data
Yes1st of 18
Community-reportedverified 2026-01-15
Zero retention
No10th of 12
Inferredverified 2026-01-15
$/M cache read
$0.028 /M tok1st of 5
Inferredverified 2026-01-15
Context window
128,000 tokens10th of 10
Inferredverified 2026-01-15
Image input
No16th of 18
Inferredverified 2026-01-15

When to use DeepSeek

DeepSeek is the right call in these situations, each one drawn from a field we actually record:

  • High-volume classification, extraction and summarisation.
  • Cost-sensitive workloads where a 35x output-price gap against a volume-tier frontier model outweighs a small quality gap.
  • Self-hosting a capable model under a genuinely permissive licence.
  • $/M output is your binding constraint. DeepSeek records $0.42 /M tok, the best figure in the LLM APIs roster. We define that field as pay-as-you-go list price per million output tokens for the named flagship model, including reasoning or thinking tokens where those are billed as output.
  • $/M input is your binding constraint. DeepSeek records $0.28 /M tok, the best figure in the LLM APIs roster. We define that field as pay-as-you-go list price per million input tokens for the named flagship model, uncached, at the standard context tier.
  • Trains on your data is your binding constraint. DeepSeek records Yes, which only 1 of the 18 tools in the LLM APIs roster do.
  • $/M cache read is your binding constraint. DeepSeek records $0.028 /M tok, the best figure in the LLM APIs roster.

When not to use DeepSeek

Reach for something else when any of the following is a requirement rather than a nice-to-have:

  • Pinnable versions. DeepSeek records No on the LLM APIs table. OpenAI records Yes on the same field. If that is a hard requirement rather than a preference, start elsewhere.
  • Zero retention. DeepSeek records No on the LLM APIs table. Amazon Bedrock records Yes on the same field. If that is a hard requirement rather than a preference, start elsewhere.
  • Context window. DeepSeek records 128,000 tokens, the weakest of the 10 tools in the LLM APIs roster. Anthropic records 1,000,000 tokens on the same field. If that is a hard requirement rather than a preference, start elsewhere.
  • Image input. DeepSeek records No on the LLM APIs table. Anthropic records Yes on the same field. If that is a hard requirement rather than a preference, start elsewhere.
  • Batch discount. DeepSeek records 0 %, 8th of 10 in the LLM APIs roster. Anthropic records 50 % on the same field. If that is a hard requirement rather than a preference, start elsewhere.

We publish this block because a comparison that only lists what a tool is good at is marketing. Every figure above sits on the same page as its source, and the field definitions are on the category tables if you want to check how we measured them.

Tools compared alongside DeepSeek

Everything below shares at least one comparison category with DeepSeek, ordered by how much overlap there is. For the reasoning on each — which fields it wins, which it loses — see DeepSeek alternatives.
Alibaba's model family — huge open-weight range, closed flagship
Wafer-scale inference — the fastest tokens per second available
Enterprise-focused models built for RAG and private deployment
Custom LPU silicon serving open-weight models at extreme speed

Recent DeepSeek changes

  • 2025-09-29 · pricingDeepSeek V3.2 cuts API prices by more than halfSparse attention in V3.2-Exp brought long-context serving costs down and DeepSeek passed it straight through, landing input around $0.28/M and output around $0.42/M. Every budget provider on this page re-priced within the quarter.
Full DeepSeek changelog

Sources and gaps

What we don't know. 4 of the 27 fields we track for DeepSeek are still blank: Max output, TTFT p50, Output tok/s and Notice period. Those render as dashes rather than as zeroes or assumptions, because an empty cell and a bad cell are not the same thing and only one of them is honest. If you know any of these figures and can point at a document, tell us.

Every figure on this page traces back to a document you can open. Where a vendor claims a number we could not reproduce, the cell says so.