Run Kimi K3: Moonshot's 2.8T Open MoE for Agents and Coding
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Kimi K3 is Moonshot AI’s new open-weight model. It is a native multimodal agentic model with 2.8T total parameters, 104B of them active per token. It has a 1,048,576-token context window and uses a new attention stack built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Moonshot calls it “the world’s first open 3T-class model.” For developers, this means you can download weights for a frontier-scale model built for long coding sessions and tool-heavy agent work, and serve them yourself. You can also call it through an OpenAI-compatible API.
Key specs
| Architecture | Mixture-of-Experts, 93 layers (1 dense) |
| Total / active parameters | 2.8T / 104B |
| Experts | 896 experts, 16 selected per token, 2 shared experts |
| Attention | 69 KDA layers + 24 Gated MLA layers |
| Context length | 1,048,576 tokens |
| Vocabulary | 160K |
| Vision encoder | MoonViT-V2, 401M parameters |
| Quantization | MXFP4 weights / MXFP8 activations, quantization-aware training from SFT onward |
| License | Kimi K3 License (custom) |
According to Moonshot, the Stable LatentMoE design gives “an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.”
The card’s benchmark table compares K3 with Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5 and GLM-5.2. A few results that help you decide whether to use it:
- SWE-Marathon: 42.0, the highest in the table (Opus 4.8: 40.0, GPT-5.6 Sol: 39.0). This row has the most caveats on the card. Moonshot ran it on an H20-calibrated branch of the official tasks, before the final v1.1 release. The card also notes that Claude Fable 5 hit fallbacks on 35% of the tasks, “which may have negatively impacted its measured performance.”
- MCPMark-Verified: 94.5 (GPT-5.6 Sol and GPT-5.5: 92.9).
- BrowseComp: 91.2 with context compaction at 300K tokens, and 90.4 using the full 1M window with no context management. The competitor scores in this row are cited from Anthropic’s and OpenAI’s own announcements, so the setups aren’t identical.
- OmniDocBench: 91.1, the highest in its row.
- Where it trails: FrontierSWE 81.2 vs. 86.6 for Claude Fable 5, OSWorld 2.0 58.3 vs. 66.1, CritPt 23.4 vs. 32.3 for GPT-5.6 Sol.
All K3 scores were run at reasoning_effort “max” and temperature 1.0. Many of the coding scores also use Moonshot’s own Kimi Code harness, while competitors use other harnesses. Read the footnotes before you treat any row as a head-to-head result.
Install
The card’s usage example talks to the model through the openai Python client. You can point that client at either the hosted API or your own server.
pip install openaiFor the hosted option, sign up at https://platform.kimi.ai and select kimi-k3. The card says the platform provides OpenAI- and Anthropic-compatible APIs. It does not give the base URL, so copy that from the platform docs.
To self-host, the card recommends three inference engines and links a recipe for each:
- vLLM: https://recipes.vllm.ai/moonshotai/Kimi-K3
- SGLang: https://docs.sglang.io/cookbook/autoregressive/Moonshotai/Kimi-K3
- TokenSpeed: https://lightseek.org/tokenspeed/recipes/models#kimi-k3
The card does not include launch commands or GPU counts, so follow the recipe for your engine.
Run it
Every example below reads the endpoint, key and model name from environment variables. kimi-k3 is the model name on platform.kimi.ai. A self-hosted server will usually expose a different name, such as the HF repo id or whatever you set as the served model name. Set KIMI_MODEL to match your server.
import osimport openai
client = openai.OpenAI( base_url=os.environ["KIMI_BASE_URL"], # from platform.kimi.ai docs, or your own server api_key=os.environ["KIMI_API_KEY"],)MODEL = os.environ.get("KIMI_MODEL", "kimi-k3") # set to your server's model name if self-hosting
response = client.chat.completions.create( model=MODEL, messages=[{"role": "user", "content": "Explain what a mixture-of-experts layer does in two sentences."}], stream=False, max_tokens=4096, reasoning_effort="high",)
msg = response.choices[0].messageprint("reasoning:", getattr(msg, "reasoning_content", None) or getattr(msg, "reasoning", None))print("answer:", msg.content)Thinking cannot be turned off. reasoning_effort accepts "low", "high" or "max", and the default is "max".
Multi-turn chat with preserved thinking
The card stresses this requirement: K3 was trained in “preserved thinking history mode.” For multi-turn conversations and tool calls, you must pass the complete assistant message returned by the API back into messages as-is. That includes reasoning_content and tool_calls, not just content.
import osimport openai
client = openai.OpenAI( base_url=os.environ["KIMI_BASE_URL"], api_key=os.environ["KIMI_API_KEY"],)MODEL = os.environ.get("KIMI_MODEL", "kimi-k3")
def ask(messages): response = client.chat.completions.create( model=MODEL, messages=messages, stream=False, max_tokens=4096, reasoning_effort="max", ) msg = response.choices[0].message # Pass the assistant message back unchanged (all returned fields, # including reasoning_content and tool_calls if present) messages.append(msg.model_dump(exclude_none=True)) return msg.content
messages = [{"role": "user", "content": "Pick five random numbers privately, then tell me only the first three."}]print(ask(messages))
messages.append({"role": "user", "content": "What were the other two numbers you had in mind?"})print(ask(messages))model_dump(exclude_none=True) serializes every field the client holds, so you don’t have to rebuild the message by hand. That only works if two things hold. First, your openai client version must keep non-standard response fields such as reasoning_content. Second, the dump must not add client-side fields the server never sent, such as annotations, audio or function_call, because the server might reject them on the next request. Print the dumped dict once to confirm the reasoning is in it and that it contains no extra keys. The card points to the Kimi K3 Quickstart (https://platform.kimi.ai/docs/guide/kimi-k3-quickstart) for tool choice, structured output and vision input. Those request formats are not on the card, so they are not shown here.
Batch classification at low effort
Every request uses reasoning, so on a high-volume task the effort setting is your main control over cost and latency. Here is a support-ticket triage loop at "low" effort:
import osimport openai
client = openai.OpenAI( base_url=os.environ["KIMI_BASE_URL"], api_key=os.environ["KIMI_API_KEY"],)MODEL = os.environ.get("KIMI_MODEL", "kimi-k3")
LABELS = ["billing", "bug", "feature_request", "account", "other"]
tickets = [ "I was charged twice for my subscription this month.", "The export button crashes the app on Safari.", "Could you add dark mode to the dashboard?", "I can't reset my password, the email never arrives.",]
def classify(text): response = client.chat.completions.create( model=MODEL, messages=[ { "role": "user", "content": ( f"Classify this support ticket into exactly one of {LABELS}. " f"Reply with the label only.\n\nTicket: {text}" ), } ], stream=False, max_tokens=1024, reasoning_effort="low", ) content = response.choices[0].message.content if content is None: return "<no answer>" # e.g. reasoning used up max_tokens return content.strip()
for t in tickets: print(f"{classify(t):16} | {t}")The card doesn’t say how max_tokens interacts with reasoning tokens. We are assuming it covers both, which is common for reasoning models, so don’t set it to a handful of tokens just because the label is short. Check the platform docs to confirm.
Long-document Q&A in one prompt
The context window is 1,048,576 tokens, so a long document can go straight into the prompt. The card lists an AA-LCR score of 74.7, the highest in its table, cited from Artificial Analysis. Test for yourself whether this approach can replace chunking or retrieval for your workload. The card doesn’t make that claim.
import osimport sysimport openai
client = openai.OpenAI( base_url=os.environ["KIMI_BASE_URL"], api_key=os.environ["KIMI_API_KEY"],)MODEL = os.environ.get("KIMI_MODEL", "kimi-k3")# 1,048,576 per the card; a self-hosted server may be configured lowerMAX_CONTEXT = int(os.environ.get("KIMI_MAX_CONTEXT", "1048576"))MAX_TOKENS = 8192
path, question = sys.argv[1], sys.argv[2]with open(path, encoding="utf-8") as f: document = f.read()
prompt = ( "Answer the question using only the document below. " "Quote the passages you rely on.\n\n" f"<document>\n{document}\n</document>\n\nQuestion: {question}")
# Rough estimate (~4 characters per token); not the model's real tokenizerest_tokens = len(prompt) // 4if est_tokens + MAX_TOKENS > MAX_CONTEXT: sys.exit(f"Prompt is ~{est_tokens} tokens; with max_tokens={MAX_TOKENS} " f"it likely exceeds the {MAX_CONTEXT}-token context.")
response = client.chat.completions.create( model=MODEL, messages=[{"role": "user", "content": prompt}], stream=False, max_tokens=MAX_TOKENS, reasoning_effort="high",)print(response.choices[0].message.content)Run it with python long_doc_qa.py contract.txt "What are the termination conditions?". The length check is only an estimate. The server will still reject a prompt that really is too long, and if it returns an error, check its configured max model length.
Gotchas
- Thinking is always on. Every call produces reasoning tokens. The default effort is
"max", so set"low"for simple tasks. - Pass back the full assistant message. Multi-turn and tool-use quality depends on sending back the complete assistant message as-is, including
reasoning_contentandtool_calls. - Field name mismatch. The card’s prose calls the field
reasoning_content, but its own example printsmessage.reasoning. The code above checks both. - Modality claims differ. The summary table lists “Text, Image,” while the introduction says K3 understands text, images and video. The card includes video benchmarks (Video-MME, MMVU), but check what your serving engine supports.
- Model name differs by endpoint.
kimi-k3is the hosted platform’s name. Self-hosted servers usually use another name, so setKIMI_MODEL. - Size. If all 2.8T parameters were stored at 4 bits, the weights alone would take roughly 1.5 TB once MXFP4 block scales are included. That is our own arithmetic, not a figure from the card. The card doesn’t say whether every component uses MXFP4, including the vision encoder, embeddings and attention layers, so treat 1.5 TB as a lower bound. The card does not list hardware requirements. MXFP4/MXFP8 is described as chosen “for broad hardware compatibility,” but this is still a cluster-scale model.
- Context compaction. On BrowseComp, K3 scored 91.2 with context compaction triggered at 300K tokens and 90.4 using the full 1M window with no context management. That is one benchmark, not general guidance, but it is worth testing compaction in long agent runs.
- Harness-dependent scores. Several coding results come from the Kimi Code harness. On Kimi Code Bench 2.0, for example, K3 scores 72.9 with Kimi Code and 73.7 with Claude Code. Expect your numbers to shift with your agent framework.
- Custom license. Weights and code are released under the Kimi K3 License, which the card does not summarize. Read the
LICENSEfile before any commercial deployment.
When to pick it
Pick Kimi K3 if you need open weights you control for long-horizon coding agents, tool-calling over MCP, deep research or document-heavy vision work, and you either have serious GPU capacity or are happy to use Moonshot’s hosted API. Its scores on SWE-Marathon, MCPMark-Verified, BrowseComp and OmniDocBench are at or near the top of the card’s comparison table, though the SWE-Marathon and BrowseComp rows come with the setup caveats noted above. Moonshot recommends pairing it with Kimi Code CLI as the agent framework.
Skip it if you need something that runs on a single GPU or a workstation. Think twice if your workload looks like OSWorld 2.0 or CritPt. On the card’s own numbers, K3 trails the top closed models there (Claude Fable 5 and GPT-5.6 Sol on both, and GPT-5.5 on CritPt), though it beats Claude Opus 4.8 on both and GPT-5.5 on OSWorld 2.0. For cheap, low-latency calls where always-on reasoning is just overhead, a smaller model will usually serve you better. Finally, if your legal team needs a standard OSI license, check the Kimi K3 License first.
Related
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
GLM-5.3-Flash: What You Can Run With Z.ai's 18B-Active Multimodal Model
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL