Skip to main content
The Human Bit

Moonshot AI · Kimi

Kimi K3

Kimi's open-weight multimodal flagship for long-horizon coding, knowledge work and tool-using tasks, with a one-million-token context window.

The Human Bit position

The weights are available, so move from launch watching to matched evaluation. Test task quality, total reasoning volume, deployment complexity and whether an open stack creates enough control or value to justify operating it.

Use it for

  • Large document and knowledge-work packs
  • Long-running coding and terminal work
  • Tasks combining visual evidence with code or tools
Leave it alone when

Do not use it as a quick-answer default, rely on its web search in production or give it broad authority when an unexpected decision would matter.

Effort without the jargon

Low

The task is simple enough that a long reasoning pass would mostly buy tokens.

High

The work needs real analysis but does not justify the deepest and most expensive pass.

Max

The task is long-horizon or difficult enough to justify the deepest reasoning pass and its token cost.

Advanced recordCost, limits, privacy and operations

K3 is now genuinely open-weight, but a 2.8-trillion-parameter model is not a simple local install. Compare hosted use, certified providers and self-deployment against the task, hardware, governance and review burden.

API cost
Officially disclosedKimi uses flat pay-as-you-go token pricing with separate cache-hit, cache-miss and output rates and no context-length tiering; check the live table for current amounts.
Context window
Officially disclosed1,000,000 tokens
Maximum output
Officially disclosed1,048,576 completion tokens (131,072 default)
Knowledge cutoff
Not disclosedNot disclosed in the cited record
Recorded inputs
Text · Images · Video
Recorded output
Text · Structured JSON

Data handling is a product decision

Not disclosedThe cited model and pricing records do not state a portable retention or training rule. Confirm Kimi API and product terms before sending sensitive material.

Operational constraints

  • Thinking is always on and cannot be disabled; reasoning effort selects low, high or max, and defaults to max.
  • Sampling values are fixed, and complete assistant messages must be preserved across tool turns.
  • Public image URLs are unsupported; use base64 or an ms:// file ID.
  • Kimi says its web-search tool is being updated and is not recommended for production in the near term.
  • The released model is 2.8 trillion total parameters and uses MXFP4 weights, so deployment still requires serious inference infrastructure and validation.
  • Multi-turn and tool workflows require the complete reasoning history returned by the API to be preserved as documented.

Unknown means the cited official record does not disclose a safe value. It is not an estimate. Prices are provider list prices in USD where stated and can change before this record's review date.

Keep the evidence separate

Canonical provider claim

Moonshot calls K3 its most capable open-weight model, with 2.8 trillion total parameters, native multimodal capability, a one-million-token context window and low, high or max reasoning effort. Full model weights, code, licence and the technical report are now released.

Independent observation · Simon Willison

His first hands-on run confirmed working vision and valid SVG output, but the reasoning pass burned 13,241 tokens to produce 3,417 tokens of answer, making one simple test cost 25 cents. Writing on 2026-07-16 he found only one reasoning effort available, max; Kimi's quickstart now documents low, high and max, so the cost trap he measured is now partly avoidable, and his token-volume warning is the part that still holds.

Read the independent source ↗
Readiness · informational

No Human Bit result is claimed; use this as evaluation guidance only.

Human Bit test

Not scheduled

No Human Bit result is claimed for this model yet.

Sources