---
title: Meta Llama
slug: meta-llama
url: "https://toolweight.com/options/meta-llama"
homepage: "https://www.llama.com"
categories: llm-apis
last_verified: 2026-01-15
license: CC-BY-4.0
---

# Meta Llama

> Open-weight Llama models, hosted almost everywhere but Meta

Meta's contribution to this category is weights, not an endpoint. Llama models are downloadable under a community licence with usage conditions, and are served by every router and cloud on this page. Meta's own hosted API exists but has never been the primary distribution channel, which makes Llama the reference case for portability: if a host retires it, three others still serve it.

## Identity

|  |  |
| --- | --- |
| Name | Meta Llama |
| Company | Meta |
| One-liner | Open-weight Llama models, hosted almost everywhere but Meta |
| Site | https://www.llama.com |
| Docs | https://www.llama.com/docs |
| Founded | 2004 |
| Open source | Yes |
| Licence | Llama Community License |
| Brand | https://toolweight.com/vendors/Meta |
| Compared in | 1 |

## Where it is compared

### [LLM APIs](https://toolweight.com/compare/llm-apis)
Ranked **#6 of 18** on default weights.
| Field | Value | Confidence | Verified | Source | Note |
| --- | --- | --- | --- | --- | --- |
| Flagship model | Llama 4 Maverick | Inferred | 2026-01-15 | - | Meta's own hosted API has never been the main distribution channel; most consumption is via routers and clouds. |
| Context window | 1,000,000 tokens | Vendor-claimed | 2026-01-15 | https://www.llama.com/models/llama-4/ | Advertised input window; usable context depends entirely on how your host has deployed it, and few hosts serve anything close to the full figure. |
| Max output | - | Unknown | - | - | Host-dependent, there is no vendor endpoint to set one. Whatever your serving stack is configured for is the answer. |
| Image input | ● | Vendor-claimed | 2026-01-15 | https://www.llama.com/models/llama-4/ | - |
| Audio in/out | ○ | Inferred | 2026-01-15 | - | - |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://www.llama.com/llama-downloads/ | Llama Community License, permissive for most commercial use but with an acceptable-use policy and a large-scale-user clause, so not OSI open source. |
| OpenAI-compat API | ● | Inferred | 2026-01-15 | - | True of essentially every host that serves Llama, rather than a property of Meta's own endpoint. |
| TTFT p50 | - | Unknown | - | - | - |
| Output tok/s | - | Unknown | - | - | - |
| $/M input | - | Unknown | - | - | Meta's first-party API pricing is not the reference point; compare Groq, Together, Fireworks or Bedrock rates for the same weights. |
| $/M output | - | Unknown | - | - | - |
| $/M cache read | - | Unknown | - | - | - |
| Batch discount | - | Unknown | - | - | - |
| Tool use | ◐ | Community-reported | 2026-01-15 | - | The models are trained for tool calling, but reliability at depth trails the closed frontier and varies by host implementation. |
| Schema output | ◐ | Inferred | 2026-01-15 | - | Depends on the serving stack, most open-weight hosts add grammar-constrained decoding. |
| Effort control | ○ | Inferred | 2026-01-15 | - | - |
| Computer use | ○ | Inferred | 2026-01-15 | - | - |
| MCP support | ○ | Inferred | 2026-01-15 | - | - |
| Cache TTL | Host-dependent | Inferred | 2026-01-15 | - | - |
| Continuity policy | Effectively unlimited, and the reason Llama matters strategically. The weights are downloadable and mirrored, so no vendor can retire the model out from under you, if a host drops it, three others still serve it, or you run it yourself. Meta may stop releasing new generations, but nothing already published disappears. | Inferred | 2026-01-15 | - | - |
| Notice period | - | Unknown | - | - | Not applicable in the usual sense, retirement is a per-host decision, not a vendor one. |
| Pinnable versions | ● | Inferred | 2026-01-15 | - | You can pin a checkpoint hash. |
| Zero retention | ◐ | Inferred | 2026-01-15 | - | Conditional on how you run it: total if you self-host, and entirely that host's policy if you do not. Regraded from yes because the weights alone guarantee nothing, the self-host baseline row is where an unconditional yes belongs. |
| Trains on your data | ○ | Inferred | 2026-01-15 | - | - |
| AU region | ● | Inferred | 2026-01-15 | - | Deployable anywhere, including Bedrock ap-southeast-2 and your own Australian hardware. |
| API since | 2023 | Inferred | - | - | - |
| Positioning | Open-weight models you can host anywhere, forever | Inferred | - | - | - |

**Verdict.** Listed here as a model family rather than a serious first-party endpoint, nobody should be calling Meta's API when Groq, Together, Fireworks and Bedrock all serve the same weights better. Its real value is as the category's continuity insurance: a capable model that literally cannot be deprecated.

## Alternatives

- [Qwen](https://toolweight.com/options/alibaba-qwen), Alibaba's model family, huge open-weight range, closed flagship
- [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock), Multi-vendor model access inside your existing AWS account
- [Anthropic](https://toolweight.com/options/anthropic), Claude models, built around long agentic runs and tool use
- [Cerebras](https://toolweight.com/options/cerebras), Wafer-scale inference, the fastest tokens per second available
- [Cohere](https://toolweight.com/options/cohere), Enterprise-focused models built for RAG and private deployment
- [DeepSeek](https://toolweight.com/options/deepseek), Frontier-adjacent models at a small fraction of Western prices
- [Fireworks AI](https://toolweight.com/options/fireworks-ai), Fast open-weight inference with strong structured-output support
- [Google Gemini](https://toolweight.com/options/google-gemini), Gemini via AI Studio for prototyping or Vertex AI for production
- [Groq](https://toolweight.com/options/groq), Custom LPU silicon serving open-weight models at extreme speed
- [Mistral AI](https://toolweight.com/options/mistral-ai), European lab with an open-weight lineage and EU-resident hosting
- [Moonshot AI](https://toolweight.com/options/moonshot-ai), Kimi models, open-weight agentic performance at low cost
- [OpenAI](https://toolweight.com/options/openai), GPT models plus audio, images and embeddings on one bill

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/options/meta-llama
- Machine-readable: https://toolweight.com/options/meta-llama.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
