---
title: Cohere
slug: cohere
url: "https://toolweight.com/options/cohere"
homepage: "https://cohere.com"
categories: llm-apis
last_verified: 2026-01-15
license: CC-BY-4.0
---

# Cohere

> Enterprise-focused models built for RAG and private deployment

Cohere sells to regulated enterprises: the Command family plus best-in-class rerank and embedding models, deployable into your own VPC or on-premise rather than only as a hosted endpoint. It has largely stopped competing on frontier benchmarks and instead competes on retrieval quality, deployment flexibility and contractual terms, which is a defensible position with a narrow audience.

## Identity

|  |  |
| --- | --- |
| Name | Cohere |
| Company | Cohere |
| One-liner | Enterprise-focused models built for RAG and private deployment |
| Site | https://cohere.com |
| Docs | https://docs.cohere.com |
| Founded | 2019 |
| Funding | Late-stage private |
| Open source | No |
| Licence | CC-BY-NC (research weights) |
| Brand | https://toolweight.com/vendors/Cohere |
| Compared in | 1 |

## Where it is compared

### [LLM APIs](https://toolweight.com/compare/llm-apis)
Ranked **#11 of 18** on default weights.
| Field | Value | Confidence | Verified | Source | Note |
| --- | --- | --- | --- | --- | --- |
| Flagship model | Command A (command-a-03-2025) | Inferred | 2026-01-15 | - | - |
| Context window | 256,000 tokens | Inferred | 2026-01-15 | - | - |
| Max output | - | Unknown | - | - | - |
| Image input | ◐ | Inferred | 2026-01-15 | - | - |
| Audio in/out | ○ | Inferred | 2026-01-15 | - | - |
| Open weights | ◐ | Inferred | 2026-01-15 | - | Weights are published for research under a non-commercial licence; commercial use requires a Cohere agreement. |
| OpenAI-compat API | ◐ | Inferred | 2026-01-15 | - | Cohere's native API is its own shape; a compatibility layer exists but is not the documented path. |
| TTFT p50 | - | Unknown | - | - | - |
| Output tok/s | - | Unknown | - | - | - |
| $/M input | $2.5 /M tok | Inferred | 2026-01-15 | https://cohere.com/pricing | - |
| $/M output | $10 /M tok | Inferred | 2026-01-15 | - | - |
| $/M cache read | - | Unknown | - | - | - |
| Batch discount | - | Unknown | - | - | - |
| Tool use | ● | Vendor-claimed | 2026-01-15 | https://docs.cohere.com/docs/tool-use | - |
| Schema output | ● | Inferred | 2026-01-15 | - | - |
| Effort control | ○ | Inferred | 2026-01-15 | - | - |
| Computer use | ○ | Inferred | 2026-01-15 | - | - |
| MCP support | ○ | Inferred | 2026-01-15 | - | - |
| Cache TTL | - | Unknown | - | - | - |
| Continuity policy | Dated model names are the default (command-a-03-2025), and Cohere publishes deprecation notices with replacement guidance. The strongest guarantee is deployment-shaped rather than policy-shaped: a private VPC or on-premise deployment continues running whatever version you licensed regardless of what the hosted platform does. | Inferred | 2026-01-15 | - | - |
| Notice period | - | Unknown | - | - | - |
| Pinnable versions | ● | Vendor-claimed | 2026-01-15 | https://docs.cohere.com/docs/models | - |
| Zero retention | ◐ | Inferred | 2026-01-15 | - | Private deployment is the primary answer, and enterprise terms cover the hosted platform, both of which are 'available on request or on a higher tier', which is what this column calls partial. Regraded down from yes: a deployment mode you have to buy is not the same as retention being off by default. |
| Trains on your data | ○ | Inferred | 2026-01-15 | - | Not on the paid platform; the free trial tier is a different agreement. |
| AU region | ◐ | Inferred | 2026-01-15 | - | Deployable into any cloud region you control, including AU; the hosted endpoint is not AU-resident. |
| API since | 2021 | Inferred | - | - | - |
| Positioning | Enterprise RAG models you can deploy inside your own network | Inferred | - | - | - |

**Verdict.** Has sensibly stopped chasing the frontier and now competes where it can win: retrieval quality, private deployment and enterprise contracts. Its rerank and embedding models remain best-in-class and are the more common reason to be a customer. As a general-purpose chat API it is priced like a frontier lab without matching one.

## Alternatives

- [Qwen](https://toolweight.com/options/alibaba-qwen), Alibaba's model family, huge open-weight range, closed flagship
- [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock), Multi-vendor model access inside your existing AWS account
- [Anthropic](https://toolweight.com/options/anthropic), Claude models, built around long agentic runs and tool use
- [Cerebras](https://toolweight.com/options/cerebras), Wafer-scale inference, the fastest tokens per second available
- [DeepSeek](https://toolweight.com/options/deepseek), Frontier-adjacent models at a small fraction of Western prices
- [Fireworks AI](https://toolweight.com/options/fireworks-ai), Fast open-weight inference with strong structured-output support
- [Google Gemini](https://toolweight.com/options/google-gemini), Gemini via AI Studio for prototyping or Vertex AI for production
- [Groq](https://toolweight.com/options/groq), Custom LPU silicon serving open-weight models at extreme speed
- [Meta Llama](https://toolweight.com/options/meta-llama), Open-weight Llama models, hosted almost everywhere but Meta
- [Mistral AI](https://toolweight.com/options/mistral-ai), European lab with an open-weight lineage and EU-resident hosting
- [Moonshot AI](https://toolweight.com/options/moonshot-ai), Kimi models, open-weight agentic performance at low cost
- [OpenAI](https://toolweight.com/options/openai), GPT models plus audio, images and embeddings on one bill

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/options/cohere
- Machine-readable: https://toolweight.com/options/cohere.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
