Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 6 min read

Run MiMo-V2.6-Distill-Qwen-9B for Coding and Agent Tasks

Xiaomi's 9B SFT model, fine-tuned from Qwen3.5-9B, scores higher than its base on SWE Pro, Terminal Bench and Toolathlon. How to serve it with SGLang.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

MiMo-V2.6-Distill-Qwen-9B is a 9B agentic model from Xiaomi MiMo. Xiaomi built it with supervised fine-tuning of Qwen3.5-9B on data generated by MiMo. It covers coding, general-purpose agent tasks, visual coding and cybersecurity. The main reason to look at it is the gap over its own base model on agent benchmarks: for example, 5.0 → 30.3 on AutomationBench and 32.0 → 44.6 on SWE Pro. Keep in mind that it is a plain SFT checkpoint. Xiaomi says it released the model “as a starting point for open research in agentic reinforcement learning,” and it has not been through RL.

Key specs

  • Base model: Qwen/Qwen3.5-9B (finetune)
  • Size: 9B parameters
  • Training: SFT only, on a weighted mixture of 77.4B total tokens (27.2B loss-bearing)
  • Data mix by token share: Code 29.9%, General 28.5%, Visual 27.4%, Cyber 14.2%
  • License: MIT
  • Serving path on the card: SGLang with --reasoning-parser mimo. The repo includes the tokenizer and the MiMo v2.6 chat template.

Results for the SFT checkpoint compared with the base model, from the MiMo-V2.6 technical report:

Benchmark Metric Qwen3.5-9B MiMo-V2.6-Distill-Qwen-9B
SWE Verified avg@3 60.0 61.1
SWE Pro avg@3 32.0 44.6
Terminal Bench 2.1 avg@1 27.0 37.1
Toolathlon-Verified avg@1 25.9 35.2
AutomationBench v1.0.6 avg@1 5.0 30.3
OfficeQA avg@1 9.0 19.5
JobBench avg@1 2.6 18.3

The card also lists four “MiMo … (mini)” results: Code, Cyber, General and Visual Coding. Those come from internal evaluation sets, so you can’t reproduce them, and I left them out of the table above. On SWE Verified the gain over the base is small (60.0 → 61.1). The bigger improvements are on SWE Pro and the general agent benchmarks. The card compares the model only against Qwen3.5-9B and has no numbers against other models.

Install

The card’s only serving path is SGLang. You need a recent SGLang build with Qwen3.5 support. Follow SGLang’s install guide for your platform. On the client side, install the openai Python package, because SGLang exposes an OpenAI-compatible endpoint.

Run it

Start the server with the command from the card, which includes the --reasoning-parser mimo flag:

Terminal window
sglang serve \
--model-path XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B \
--reasoning-parser mimo \
--host 0.0.0.0 \
--port 30000

Then query it with thinking turned on explicitly:

quickstart.py
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:30000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B",
messages=[
{"role": "user", "content": "What is 15% of 240?"}
],
max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
message = response.choices[0].message
print("Thinking:", getattr(message, "reasoning_content", "") or "")
print("Answer:", message.content or "")

The card’s example reads the reasoning trace from reasoning_content and the final answer from content. Both examples below use this same call.

Fix a bug from a traceback

Code has the largest share of total training tokens, and SWE Pro shows the biggest gain among the public code benchmarks (32.0 → 44.6). Note that the gain on SWE Verified is much smaller (60.0 → 61.1), so don’t expect a big jump on every kind of bug-fixing task. This script sends a failing function and its traceback, then prints a proposed fix. The reasoning trace goes to a separate file, so you can check how the model got to its answer without cluttering the output.

fix_bug.py
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:30000/v1",
api_key="EMPTY",
)
source = '''
def average(values):
return sum(values) / len(values)
print(average([]))
'''
traceback_text = '''
Traceback (most recent call last):
File "stats.py", line 4, in <module>
print(average([]))
File "stats.py", line 2, in average
return sum(values) / len(values)
ZeroDivisionError: division by zero
'''
prompt = (
"Fix the bug in this Python code. Return the corrected function in a "
"single code block, followed by one sentence explaining the change.\n\n"
f"Code:\n{source}\n\nTraceback:\n{traceback_text}"
)
response = client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B",
messages=[{"role": "user", "content": prompt}],
max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
message = response.choices[0].message
with open("fix_bug_reasoning.txt", "w") as f:
f.write(getattr(message, "reasoning_content", "") or "")
print(message.content or "")

Triage security findings in batch

Cybersecurity data makes up 14.2% of the training tokens. The only cyber result on the card is from an internal set (5.7 → 31.3), so test the model on your own data before you rely on it. This example sends a list of scanner findings to the server in parallel, using a thread pool on the client side, and asks for a severity label and a short justification for each one.

triage.py
from concurrent.futures import ThreadPoolExecutor
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:30000/v1",
api_key="EMPTY",
)
findings = [
"Flask app runs with debug=True in production config.",
"SQL query built with f-string: f\"SELECT * FROM users WHERE id={user_id}\"",
"Dependency 'requests' pinned to a version two minor releases behind latest.",
"AWS secret key committed in config/settings.py.",
]
def triage(finding: str) -> dict:
prompt = (
"You are reviewing a security scanner finding. Reply with exactly two "
"lines:\nSEVERITY: one of critical, high, medium, low\n"
"REASON: one sentence\n\n"
f"Finding: {finding}"
)
response = client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B",
messages=[{"role": "user", "content": prompt}],
max_tokens=2048,
extra_body={"chat_template_kwargs": {"enable_thinking": True}},
)
message = response.choices[0].message
return {"finding": finding, "verdict": (message.content or "").strip()}
with ThreadPoolExecutor(max_workers=4) as pool:
results = list(pool.map(triage, findings))
for r in results:
print(r["finding"])
print(r["verdict"])
print("-" * 60)

Parse the SEVERITY: line yourself and treat anything that doesn’t match the expected format as “needs review”. The card says nothing about structured output, so don’t assume the format will always come back exactly as requested.

Gotchas

  • This is an SFT checkpoint, not a finished RL model. Xiaomi presents it as a starting point for agentic RL research, and it has not had an RL stage.
  • Use the MiMo reasoning parser. The card serves the model with --reasoning-parser mimo and reads the thinking from reasoning_content. The card doesn’t explain what happens without the parser. My guess is that the reasoning would not come back separately in reasoning_content, so keep the flag.
  • Thinking is enabled through the chat template. The card passes extra_body={"chat_template_kwargs": {"enable_thinking": True}}. Use the chat template that ships with the checkpoint (MiMo v2.6) rather than a generic Qwen template.
  • Watch max_tokens. The card’s example uses 2048. If an answer comes back empty or cut off, raise the limit before deciding the model failed.
  • Vision isn’t documented yet. Visual data makes up 27.4% of training, and the card reports a visual coding result. But the quickstart only covers text generation and doesn’t show how to send images. Wait for official instructions before you build on image input.
  • Tool calling isn’t documented either. The model is described as agentic and tagged tool-use, but the card doesn’t show a tool-calling configuration for SGLang. Check the SGLang docs or the technical report before you wire it into an agent framework.
  • You need a new SGLang build. The card asks for a recent build with Qwen3.5 support, so older installs may not load the model.
  • Hardware isn’t specified. The card gives no memory or GPU requirements.
  • Several benchmarks are internal. Every “MiMo … (mini)” score comes from an unpublished set.

When to pick it

Pick it if you want a 9B open model you can self-host for coding and terminal or automation agent work, and you’re currently on Qwen3.5-9B. The card shows higher scores than that base on every benchmark it reports, with large gains across the general agent benchmarks (AutomationBench, JobBench, OfficeQA, Terminal Bench 2.1, Toolathlon-Verified) and SWE Pro. The MIT license and the “start RL from here” framing also make it a reasonable base if you’re doing your own agentic RL or further fine-tuning.

Skip it, at least for now:

  • If you need image input in production. The card doesn’t show how to do it.
  • If you need documented tool-calling setup out of the box. The card doesn’t provide it.
  • If you mainly care about SWE Verified-style bug fixing. The improvement over the base is about one point there, so switching may not be worth the effort.
  • If you need a comparison with other models in the 9B class. The card only compares against Qwen3.5-9B, so run your own evaluation.

Related