Self-hosted (vLLM)

Self-hosted (vLLM) is run open weights on your own GPUs behind an OpenAI-shaped API. We compare it in LLM APIs. It is one of only 2 of 12 tools in its roster with Zero retention. The trade-off is being 8th of 10 on Batch discount (0 %).

What is Self-hosted (vLLM)?

Self-hosted (vLLM) — Run open weights on your own GPUs behind an OpenAI-shaped API. The baseline every hosted option is measured against. vLLM serves open-weight models on your own hardware with an OpenAI-compatible endpoint, continuous batching, paged attention and prefix caching. You trade per-token billing for GPU-hour billing and an operational burden, and you get the only genuinely permanent answer to deprecation: the weights are on your disk and nobody can retire them. Founded 2023. Open source under Apache-2.0. We carry it as the do-it-yourself baseline, not as a vendor.

We track it in 1 comparison — LLM APIs — so every claim below is a cell in a table you can open and check rather than an impression. Across those rosters it sits against 17 other tools, and what follows is where it visibly separates from them.

Where it wins.

  • Zero retention — Yes, which only 2 of 12 tools here manage.
  • AU region — Yes, which only 4 of 17 tools here manage.

Each of those is ranked against the whole roster on its category page, not against a hand-picked subset, so a first place here means first of everything we list.

Where it gives ground.

  • Batch discount — 0 %, 8th of 10. Anthropic records 50 %.
  • Effort control — No. Anthropic records Yes.

None of these disqualify it on their own. They are the fields to check against your own requirements before you commit, because they are the ones where a competitor genuinely does better.

Provenance. 19 of 27 tracked fields carry a value for Self-hosted (vLLM), and 3 of those cite a document you can open. Last verified 2026-01-15. Every figure keeps its own provenance — measured by us, claimed by the vendor, inferred, or community-reported — and we would rather print a dash than a guess.

Its nearest neighbour in our data is Qwen. Self-hosted (vLLM) is ahead on Schema output (Yes against Partial) and Open weights (Yes against Partial). Qwen takes Batch discount (50 % against 0 %) and Trains on your data (Partial against No). That pattern repeats across the rest of the roster — see Self-hosted (vLLM) alternatives for the other rivals, each compared the same way.

At a glance
Founded
2023
Licence
Apache-2.0 (open source)
Fields we track
19 of 27
Last verified
2026-01-15

Self-hosted (vLLM) in LLM APIs

Ranked against 18 tools across 27 sourced fields. Open the full LLM APIs table.

The reference point rather than a recommendation. Self-hosting only beats a hosted API on cost at genuinely high, steady utilisation, and it costs you the agentic features — effort control, computer use, MCP — that make the frontier APIs worth their price. Choose it for sovereignty or permanence, not to save money.

Where it lands in this roster
Zero retention
Yes1st of 12
Inferredverified 2026-01-15
Batch discount
0 %8th of 10
Inferred
AU region
Yes1st of 17
Inferredverified 2026-01-15
Schema output
Yes1st of 18
Vendor-claimedverified 2026-01-15source
Pinnable versions
Yes1st of 18
Inferredverified 2026-01-15
Trains on your data
No7th of 18
Inferredverified 2026-01-15
Effort control
No12th of 18
Inferredverified 2026-01-15
OpenAI-compat API
Yes1st of 18
Vendor-claimedverified 2026-01-15source

When to use Self-hosted (vLLM)

Self-hosted (vLLM) is the right call in these situations, each one drawn from a field we actually record:

  • Data that legally cannot leave your infrastructure.
  • Sustained high-utilisation inference where GPU-hours beat per-token billing.
  • Guaranteeing a model version survives indefinitely.
  • Zero retention is your binding constraint. Self-hosted (vLLM) records Yes, which only 2 of the 12 tools in the LLM APIs roster do. We define that field as whether prompts and completions can be excluded from all storage, including abuse-monitoring logs;.
  • AU region is your binding constraint. Self-hosted (vLLM) records Yes, which only 4 of the 17 tools in the LLM APIs roster do. We define that field as whether inference can be pinned to Australian infrastructure;.

When not to use Self-hosted (vLLM)

Reach for something else when any of the following is a requirement rather than a nice-to-have:

  • Batch discount. Self-hosted (vLLM) records 0 %, 8th of 10 in the LLM APIs roster. Anthropic records 50 % on the same field. If that is a hard requirement rather than a preference, start elsewhere.
  • Effort control. Self-hosted (vLLM) records No on the LLM APIs table. Anthropic records Yes on the same field. If that is a hard requirement rather than a preference, start elsewhere.

We publish this block because a comparison that only lists what a tool is good at is marketing. Every figure above sits on the same page as its source, and the field definitions are on the category tables if you want to check how we measured them.

Tools compared alongside Self-hosted (vLLM)

Everything below shares at least one comparison category with Self-hosted (vLLM), ordered by how much overlap there is. For the reasoning on each — which fields it wins, which it loses — see Self-hosted (vLLM) alternatives.
Alibaba's model family — huge open-weight range, closed flagship
Wafer-scale inference — the fastest tokens per second available
Enterprise-focused models built for RAG and private deployment

Recent Self-hosted (vLLM) changes

We have not logged a dated change for Self-hosted (vLLM) yet. The timeline fills in as pricing moves, features ship and things get deprecated.

Sources and gaps

What we don't know. 8 of the 27 fields we track for Self-hosted (vLLM) are still blank: Context window, Max output, TTFT p50, Output tok/s, $/M input, $/M output, $/M cache read and Notice period. Those render as dashes rather than as zeroes or assumptions, because an empty cell and a bad cell are not the same thing and only one of them is honest. If you know any of these figures and can point at a document, tell us.

Every figure on this page traces back to a document you can open. Where a vendor claims a number we could not reproduce, the cell says so.