Moonshot AI · Kimi
Kimi K3
Kimi's open-weight multimodal flagship for long-horizon coding, knowledge work and tool-using tasks, with a one-million-token context window.
The weights are available, so move from launch watching to matched evaluation. Test task quality, total reasoning volume, deployment complexity and whether an open stack creates enough control or value to justify operating it.
Use it for
- Large document and knowledge-work packs
- Long-running coding and terminal work
- Tasks combining visual evidence with code or tools
Do not use it as a quick-answer default, rely on its web search in production or give it broad authority when an unexpected decision would matter.
Effort without the jargon
The task is simple enough that a long reasoning pass would mostly buy tokens.
The work needs real analysis but does not justify the deepest and most expensive pass.
The task is long-horizon or difficult enough to justify the deepest reasoning pass and its token cost.
Advanced recordCost, limits, privacy and operations
K3 is now genuinely open-weight, but a 2.8-trillion-parameter model is not a simple local install. Compare hosted use, certified providers and self-deployment against the task, hardware, governance and review burden.
- API cost
- Officially disclosedKimi uses flat pay-as-you-go token pricing with separate cache-hit, cache-miss and output rates and no context-length tiering; check the live table for current amounts.
- Context window
- Officially disclosed1,000,000 tokens
- Maximum output
- Officially disclosed1,048,576 completion tokens (131,072 default)
- Knowledge cutoff
- Not disclosedNot disclosed in the cited record
- Recorded inputs
- Text · Images · Video
- Recorded output
- Text · Structured JSON
Data handling is a product decision
Not disclosedThe cited model and pricing records do not state a portable retention or training rule. Confirm Kimi API and product terms before sending sensitive material.
Operational constraints
- Thinking is always on and cannot be disabled; reasoning effort selects low, high or max, and defaults to max.
- Sampling values are fixed, and complete assistant messages must be preserved across tool turns.
- Public image URLs are unsupported; use base64 or an ms:// file ID.
- Kimi says its web-search tool is being updated and is not recommended for production in the near term.
- The released model is 2.8 trillion total parameters and uses MXFP4 weights, so deployment still requires serious inference infrastructure and validation.
- Multi-turn and tool workflows require the complete reasoning history returned by the API to be preserved as documented.
Unknown means the cited official record does not disclose a safe value. It is not an estimate. Prices are provider list prices in USD where stated and can change before this record's review date.
Keep the evidence separate
Moonshot calls K3 its most capable open-weight model, with 2.8 trillion total parameters, native multimodal capability, a one-million-token context window and low, high or max reasoning effort. Full model weights, code, licence and the technical report are now released.
His first hands-on run confirmed working vision and valid SVG output, but the reasoning pass burned 13,241 tokens to produce 3,417 tokens of answer, making one simple test cost 25 cents. Writing on 2026-07-16 he found only one reasoning effort available, max; Kimi's quickstart now documents low, high and max, so the cost trap he measured is now partly avoidable, and his token-volume warning is the part that still holds.
Read the independent source ↗No Human Bit result is claimed; use this as evaluation guidance only.
Not scheduled
No Human Bit result is claimed for this model yet.