---
title: Moonshot AI
slug: moonshot-ai
url: "https://toolweight.com/options/moonshot-ai"
homepage: "https://www.moonshot.ai"
categories: llm-apis
last_verified: 2026-01-15
license: CC-BY-4.0
---

# Moonshot AI

> Kimi models, open-weight agentic performance at low cost

Moonshot's Kimi K2 line is the open-weight model most often cited as competitive with closed frontier models on agentic and tool-use benchmarks, released under a permissive modified-MIT licence. The first-party platform is inexpensive and OpenAI-compatible; the weights are also hosted by most routers on this page, which is the practical route for anyone who cannot send data to a mainland-China endpoint.

## Identity

|  |  |
| --- | --- |
| Name | Moonshot AI |
| Company | Moonshot AI |
| One-liner | Kimi models, open-weight agentic performance at low cost |
| Site | https://www.moonshot.ai |
| Docs | https://platform.moonshot.ai/docs |
| Founded | 2023 |
| Open source | Yes |
| Licence | Modified MIT |
| Brand | https://toolweight.com/vendors/Moonshot AI |
| Compared in | 1 |

## Where it is compared

### [LLM APIs](https://toolweight.com/compare/llm-apis)
Ranked **#15 of 18** on default weights.
| Field | Value | Confidence | Verified | Source | Note |
| --- | --- | --- | --- | --- | --- |
| Flagship model | Kimi K2 Thinking | Inferred | 2026-01-15 | - | - |
| Context window | 256,000 tokens | Inferred | 2026-01-15 | - | - |
| Max output | - | Unknown | - | - | - |
| Image input | ○ | Inferred | 2026-01-15 | - | - |
| Audio in/out | ○ | Inferred | 2026-01-15 | - | - |
| Open weights | ● | Vendor-claimed | 2026-01-15 | https://github.com/MoonshotAI/Kimi-K2 | Modified MIT licence with an attribution condition for very large deployments. |
| OpenAI-compat API | ● | Vendor-claimed | 2026-01-15 | https://platform.moonshot.ai/docs/api/chat | - |
| TTFT p50 | - | Unknown | - | - | - |
| Output tok/s | - | Unknown | - | - | Slow on the first-party endpoint by most community reports; Groq and Cerebras serve the same weights far faster. |
| $/M input | $0.6 /M tok | Inferred | 2026-01-15 | - | Cache-miss rate; cache hits are roughly a quarter of this. |
| $/M output | $2.5 /M tok | Inferred | 2026-01-15 | - | - |
| $/M cache read | $0.15 /M tok | Inferred | 2026-01-15 | - | - |
| Batch discount | - | Unknown | - | - | - |
| Tool use | ● | Community-reported | 2026-01-15 | - | The open-weight model most often reported as competitive with closed frontier models on agentic benchmarks. |
| Schema output | ◐ | Inferred | 2026-01-15 | - | - |
| Effort control | ◐ | Inferred | 2026-01-15 | - | - |
| Computer use | ○ | Inferred | 2026-01-15 | - | - |
| MCP support | ○ | Inferred | 2026-01-15 | - | - |
| Cache TTL | Explicit context caching, paid per storage hour | Inferred | 2026-01-15 | - | - |
| Continuity policy | Dated model IDs (the -0905 style suffix) are used and remain addressable, which is better than DeepSeek's rolling aliases. No published deprecation policy or notice window. As with the other open-weight labs, the real guarantee is the licence, download the checkpoint and continuity is your problem, not theirs. | Inferred | 2026-01-15 | - | - |
| Notice period | - | Unknown | - | - | - |
| Pinnable versions | ● | Inferred | 2026-01-15 | - | - |
| Zero retention | ○ | Inferred | 2026-01-15 | - | - |
| Trains on your data | ◐ | Inferred | 2026-01-15 | - | - |
| AU region | ○ | Inferred | 2026-01-15 | - | - |
| API since | 2023 | Inferred | - | - | - |
| Positioning | Open-weight agentic performance at open-weight prices | Inferred | - | - | - |

**Verdict.** Kimi K2 is the strongest argument that open weights have caught up on agentic work specifically, it holds up in tool loops where other open models fall apart. Reach it through Groq, Fireworks or Together rather than the first-party endpoint, which is slower and has weaker data terms.

## Alternatives

- [Qwen](https://toolweight.com/options/alibaba-qwen), Alibaba's model family, huge open-weight range, closed flagship
- [Amazon Bedrock](https://toolweight.com/options/amazon-bedrock), Multi-vendor model access inside your existing AWS account
- [Anthropic](https://toolweight.com/options/anthropic), Claude models, built around long agentic runs and tool use
- [Cerebras](https://toolweight.com/options/cerebras), Wafer-scale inference, the fastest tokens per second available
- [Cohere](https://toolweight.com/options/cohere), Enterprise-focused models built for RAG and private deployment
- [DeepSeek](https://toolweight.com/options/deepseek), Frontier-adjacent models at a small fraction of Western prices
- [Fireworks AI](https://toolweight.com/options/fireworks-ai), Fast open-weight inference with strong structured-output support
- [Google Gemini](https://toolweight.com/options/google-gemini), Gemini via AI Studio for prototyping or Vertex AI for production
- [Groq](https://toolweight.com/options/groq), Custom LPU silicon serving open-weight models at extreme speed
- [Meta Llama](https://toolweight.com/options/meta-llama), Open-weight Llama models, hosted almost everywhere but Meta
- [Mistral AI](https://toolweight.com/options/mistral-ai), European lab with an open-weight lineage and EU-resident hosting
- [OpenAI](https://toolweight.com/options/openai), GPT models plus audio, images and embeddings on one bill

## Licence and attribution

Data from toolweight (https://toolweight.com), licensed CC-BY-4.0.

- Licence: [CC-BY-4.0](https://creativecommons.org/licenses/by/4.0/)
- Canonical HTML: https://toolweight.com/options/moonshot-ai
- Machine-readable: https://toolweight.com/options/moonshot-ai.md · https://toolweight.com/api/v1 · https://toolweight.com/mcp
- toolweight takes no affiliate revenue and sells no placements. Corrections: https://toolweight.com/suggest
