Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 8 min read

Run Kimi K3: Moonshot's 2.8T Open MoE for Agents and Coding

Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Kimi K3 is Moonshot AI’s new open-weight model. It is a native multimodal agentic model with 2.8T total parameters, 104B of them active per token. It has a 1,048,576-token context window and uses a new attention stack built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes). Moonshot calls it “the world’s first open 3T-class model.” For developers, this means you can download weights for a frontier-scale model built for long coding sessions and tool-heavy agent work, and serve them yourself. You can also call it through an OpenAI-compatible API.

Key specs

Architecture Mixture-of-Experts, 93 layers (1 dense)
Total / active parameters 2.8T / 104B
Experts 896 experts, 16 selected per token, 2 shared experts
Attention 69 KDA layers + 24 Gated MLA layers
Context length 1,048,576 tokens
Vocabulary 160K
Vision encoder MoonViT-V2, 401M parameters
Quantization MXFP4 weights / MXFP8 activations, quantization-aware training from SFT onward
License Kimi K3 License (custom)

According to Moonshot, the Stable LatentMoE design gives “an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.”

The card’s benchmark table compares K3 with Claude Fable 5, GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5 and GLM-5.2. A few results that help you decide whether to use it:

  • SWE-Marathon: 42.0, the highest in the table (Opus 4.8: 40.0, GPT-5.6 Sol: 39.0). This row has the most caveats on the card. Moonshot ran it on an H20-calibrated branch of the official tasks, before the final v1.1 release. The card also notes that Claude Fable 5 hit fallbacks on 35% of the tasks, “which may have negatively impacted its measured performance.”
  • MCPMark-Verified: 94.5 (GPT-5.6 Sol and GPT-5.5: 92.9).
  • BrowseComp: 91.2 with context compaction at 300K tokens, and 90.4 using the full 1M window with no context management. The competitor scores in this row are cited from Anthropic’s and OpenAI’s own announcements, so the setups aren’t identical.
  • OmniDocBench: 91.1, the highest in its row.
  • Where it trails: FrontierSWE 81.2 vs. 86.6 for Claude Fable 5, OSWorld 2.0 58.3 vs. 66.1, CritPt 23.4 vs. 32.3 for GPT-5.6 Sol.

All K3 scores were run at reasoning_effort “max” and temperature 1.0. Many of the coding scores also use Moonshot’s own Kimi Code harness, while competitors use other harnesses. Read the footnotes before you treat any row as a head-to-head result.

Install

The card’s usage example talks to the model through the openai Python client. You can point that client at either the hosted API or your own server.

Terminal window
pip install openai

For the hosted option, sign up at https://platform.kimi.ai and select kimi-k3. The card says the platform provides OpenAI- and Anthropic-compatible APIs. It does not give the base URL, so copy that from the platform docs.

To self-host, the card recommends three inference engines and links a recipe for each:

The card does not include launch commands or GPU counts, so follow the recipe for your engine.

Run it

Every example below reads the endpoint, key and model name from environment variables. kimi-k3 is the model name on platform.kimi.ai. A self-hosted server will usually expose a different name, such as the HF repo id or whatever you set as the served model name. Set KIMI_MODEL to match your server.

hello_k3.py
import os
import openai
client = openai.OpenAI(
base_url=os.environ["KIMI_BASE_URL"], # from platform.kimi.ai docs, or your own server
api_key=os.environ["KIMI_API_KEY"],
)
MODEL = os.environ.get("KIMI_MODEL", "kimi-k3") # set to your server's model name if self-hosting
response = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": "Explain what a mixture-of-experts layer does in two sentences."}],
stream=False,
max_tokens=4096,
reasoning_effort="high",
)
msg = response.choices[0].message
print("reasoning:", getattr(msg, "reasoning_content", None) or getattr(msg, "reasoning", None))
print("answer:", msg.content)

Thinking cannot be turned off. reasoning_effort accepts "low", "high" or "max", and the default is "max".

Multi-turn chat with preserved thinking

The card stresses this requirement: K3 was trained in “preserved thinking history mode.” For multi-turn conversations and tool calls, you must pass the complete assistant message returned by the API back into messages as-is. That includes reasoning_content and tool_calls, not just content.

multiturn.py
import os
import openai
client = openai.OpenAI(
base_url=os.environ["KIMI_BASE_URL"],
api_key=os.environ["KIMI_API_KEY"],
)
MODEL = os.environ.get("KIMI_MODEL", "kimi-k3")
def ask(messages):
response = client.chat.completions.create(
model=MODEL,
messages=messages,
stream=False,
max_tokens=4096,
reasoning_effort="max",
)
msg = response.choices[0].message
# Pass the assistant message back unchanged (all returned fields,
# including reasoning_content and tool_calls if present)
messages.append(msg.model_dump(exclude_none=True))
return msg.content
messages = [{"role": "user", "content": "Pick five random numbers privately, then tell me only the first three."}]
print(ask(messages))
messages.append({"role": "user", "content": "What were the other two numbers you had in mind?"})
print(ask(messages))

model_dump(exclude_none=True) serializes every field the client holds, so you don’t have to rebuild the message by hand. That only works if two things hold. First, your openai client version must keep non-standard response fields such as reasoning_content. Second, the dump must not add client-side fields the server never sent, such as annotations, audio or function_call, because the server might reject them on the next request. Print the dumped dict once to confirm the reasoning is in it and that it contains no extra keys. The card points to the Kimi K3 Quickstart (https://platform.kimi.ai/docs/guide/kimi-k3-quickstart) for tool choice, structured output and vision input. Those request formats are not on the card, so they are not shown here.

Batch classification at low effort

Every request uses reasoning, so on a high-volume task the effort setting is your main control over cost and latency. Here is a support-ticket triage loop at "low" effort:

batch_classify.py
import os
import openai
client = openai.OpenAI(
base_url=os.environ["KIMI_BASE_URL"],
api_key=os.environ["KIMI_API_KEY"],
)
MODEL = os.environ.get("KIMI_MODEL", "kimi-k3")
LABELS = ["billing", "bug", "feature_request", "account", "other"]
tickets = [
"I was charged twice for my subscription this month.",
"The export button crashes the app on Safari.",
"Could you add dark mode to the dashboard?",
"I can't reset my password, the email never arrives.",
]
def classify(text):
response = client.chat.completions.create(
model=MODEL,
messages=[
{
"role": "user",
"content": (
f"Classify this support ticket into exactly one of {LABELS}. "
f"Reply with the label only.\n\nTicket: {text}"
),
}
],
stream=False,
max_tokens=1024,
reasoning_effort="low",
)
content = response.choices[0].message.content
if content is None:
return "<no answer>" # e.g. reasoning used up max_tokens
return content.strip()
for t in tickets:
print(f"{classify(t):16} | {t}")

The card doesn’t say how max_tokens interacts with reasoning tokens. We are assuming it covers both, which is common for reasoning models, so don’t set it to a handful of tokens just because the label is short. Check the platform docs to confirm.

Long-document Q&A in one prompt

The context window is 1,048,576 tokens, so a long document can go straight into the prompt. The card lists an AA-LCR score of 74.7, the highest in its table, cited from Artificial Analysis. Test for yourself whether this approach can replace chunking or retrieval for your workload. The card doesn’t make that claim.

long_doc_qa.py
import os
import sys
import openai
client = openai.OpenAI(
base_url=os.environ["KIMI_BASE_URL"],
api_key=os.environ["KIMI_API_KEY"],
)
MODEL = os.environ.get("KIMI_MODEL", "kimi-k3")
# 1,048,576 per the card; a self-hosted server may be configured lower
MAX_CONTEXT = int(os.environ.get("KIMI_MAX_CONTEXT", "1048576"))
MAX_TOKENS = 8192
path, question = sys.argv[1], sys.argv[2]
with open(path, encoding="utf-8") as f:
document = f.read()
prompt = (
"Answer the question using only the document below. "
"Quote the passages you rely on.\n\n"
f"<document>\n{document}\n</document>\n\nQuestion: {question}"
)
# Rough estimate (~4 characters per token); not the model's real tokenizer
est_tokens = len(prompt) // 4
if est_tokens + MAX_TOKENS > MAX_CONTEXT:
sys.exit(f"Prompt is ~{est_tokens} tokens; with max_tokens={MAX_TOKENS} "
f"it likely exceeds the {MAX_CONTEXT}-token context.")
response = client.chat.completions.create(
model=MODEL,
messages=[{"role": "user", "content": prompt}],
stream=False,
max_tokens=MAX_TOKENS,
reasoning_effort="high",
)
print(response.choices[0].message.content)

Run it with python long_doc_qa.py contract.txt "What are the termination conditions?". The length check is only an estimate. The server will still reject a prompt that really is too long, and if it returns an error, check its configured max model length.

Gotchas

  • Thinking is always on. Every call produces reasoning tokens. The default effort is "max", so set "low" for simple tasks.
  • Pass back the full assistant message. Multi-turn and tool-use quality depends on sending back the complete assistant message as-is, including reasoning_content and tool_calls.
  • Field name mismatch. The card’s prose calls the field reasoning_content, but its own example prints message.reasoning. The code above checks both.
  • Modality claims differ. The summary table lists “Text, Image,” while the introduction says K3 understands text, images and video. The card includes video benchmarks (Video-MME, MMVU), but check what your serving engine supports.
  • Model name differs by endpoint. kimi-k3 is the hosted platform’s name. Self-hosted servers usually use another name, so set KIMI_MODEL.
  • Size. If all 2.8T parameters were stored at 4 bits, the weights alone would take roughly 1.5 TB once MXFP4 block scales are included. That is our own arithmetic, not a figure from the card. The card doesn’t say whether every component uses MXFP4, including the vision encoder, embeddings and attention layers, so treat 1.5 TB as a lower bound. The card does not list hardware requirements. MXFP4/MXFP8 is described as chosen “for broad hardware compatibility,” but this is still a cluster-scale model.
  • Context compaction. On BrowseComp, K3 scored 91.2 with context compaction triggered at 300K tokens and 90.4 using the full 1M window with no context management. That is one benchmark, not general guidance, but it is worth testing compaction in long agent runs.
  • Harness-dependent scores. Several coding results come from the Kimi Code harness. On Kimi Code Bench 2.0, for example, K3 scores 72.9 with Kimi Code and 73.7 with Claude Code. Expect your numbers to shift with your agent framework.
  • Custom license. Weights and code are released under the Kimi K3 License, which the card does not summarize. Read the LICENSE file before any commercial deployment.

When to pick it

Pick Kimi K3 if you need open weights you control for long-horizon coding agents, tool-calling over MCP, deep research or document-heavy vision work, and you either have serious GPU capacity or are happy to use Moonshot’s hosted API. Its scores on SWE-Marathon, MCPMark-Verified, BrowseComp and OmniDocBench are at or near the top of the card’s comparison table, though the SWE-Marathon and BrowseComp rows come with the setup caveats noted above. Moonshot recommends pairing it with Kimi Code CLI as the agent framework.

Skip it if you need something that runs on a single GPU or a workstation. Think twice if your workload looks like OSWorld 2.0 or CritPt. On the card’s own numbers, K3 trails the top closed models there (Claude Fable 5 and GPT-5.6 Sol on both, and GPT-5.5 on CritPt), though it beats Claude Opus 4.8 on both and GPT-5.5 on OSWorld 2.0. For cheap, low-latency calls where always-on reasoning is just overhead, a smaller model will usually serve you better. Finally, if your legal team needs a standard OSI license, check the Kimi K3 License first.

Related