Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 9 min read

Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model

Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

MiMo-V2.6-Pro-RL is the flagship checkpoint of Xiaomi’s MiMo-V2.6 series. It is a sparse Mixture-of-Experts model with 1.02T total parameters, 42B of them active per token. It takes text, image, video and audio, and its context window is 1M tokens. Most of the card is about how the model was trained. Xiaomi ran one mixed reinforcement-learning run across coding, general agents, visual tasks and cybersecurity, rather than one run per domain. A “groupwise agentic grader” then ranks passing solutions against each other instead of scoring them only pass/fail. The model is released under MIT. On several agent benchmarks it scores close to the closed models the card compares it with. On others it is well behind, including Terminal Bench 4.0 and the exploit benchmarks.

Key specs

MiMo-V2.6-Pro-RL
Architecture Sparse MoE, 1.02T total / 42B activated
Experts 384 routed, 8 activated, no shared experts
Layers 70 (60 sliding-window, 10 global attention)
Sliding window 128 tokens
Context 1M tokens
Modalities Text, image, video, audio
Vision encoder 681M-param MiMo ViT
Audio 308M AudioTokenizer + 127M audio patch encoder
Speculative decoder 5-layer MTP drafter, predicts 7 tokens per pass
License MIT

These are selected benchmark results from the card. The “Pro” column is this model, and the others are the comparison models the card lists:

Benchmark V2.6 Pro V2.6 Flash Claude Opus 5 GPT-5.6 Sol
DeepSWE v1.1 71.9 67.9 74.0 73.0
AutomationBench v1.0.6 53.1 52.3 50.3 45.8
Toolathlon-Verified 76.9 73.6 80.6 74.9
Terminal Bench 2.1 89.9 87.6 89.1 88.8
Terminal Bench 4.0 34.9 28.8 49.0 39.9
OSWorld-Verified 82.0 80.8 83.4 83.0
ExploitBench 47.9 25.3 70.0 78.5

Among the models in this table, it has the top score on AutomationBench and Terminal Bench 2.1. It trails Claude Opus 5 on DeepSWE, Toolathlon, Terminal Bench 4.0 and OSWorld-Verified. It also trails GPT-5.6 Sol on DeepSWE (71.9 vs 73.0), Terminal Bench 4.0 (34.9 vs 39.9) and OSWorld-Verified (82.0 vs 83.0). It is well behind both closed models on ExploitBench. Compared with the previous MiMo-V2.5 Pro, the gains are large: DeepSWE went from 19.0 to 71.9 and AutomationBench from 16.0 to 53.1.

Install

The card gives two serving paths and names a prebuilt Docker image for each. The vLLM image and both linked recipes are MiMo-V2.5 artifacts: the vLLM tag is mimov25-cu129, and the cookbook and recipe pages are for V2.5. The SGLang image is the generic lmsysorg/sglang:latest. Check the recipe pages for any V2.6-specific changes before you rely on them.

Terminal window
# vLLM image (from the vLLM MiMo-V2.5 recipe)
docker pull vllm/vllm-openai:mimov25-cu129
# SGLang image (named alongside the SGLang MiMo cookbook)
docker pull lmsysorg/sglang:latest

The card doesn’t include a docker run command, so it doesn’t cover GPU flags, port mapping or cache mounts. Follow the vLLM MiMo-V2.5 recipe or the SGLang MiMo cookbook to start the container. Run the serve commands below inside that container, or on a host where vLLM or SGLang is already installed.

To query the server from Python, install the OpenAI client:

Terminal window
pip install openai

Run it

The card’s vLLM command is the shorter of the two. It uses tensor parallelism of 8:

Terminal window
vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL \
--tensor-parallel-size 8 \
--trust-remote-code \
--gpu-memory-utilization 0.95 \
--max-model-len auto \
--reasoning-parser mimo \
--tool-call-parser mimo \
--enable-auto-tool-choice \
--generation-config vllm

The card gives no GPU type, memory figure or weight precision. Check that a 1.02T-parameter model fits on your 8 GPUs before you start.

“For best performance,” the card points to the SGLang MiMo cookbook. Xiaomi’s reference SGLang setup runs across 2 nodes with TP=16, DP=2, expert parallelism, DeepEP all-to-all and EAGLE speculative decoding. The full command is in the model card and the SGLang MiMo cookbook. It listens on port 30000.

The model card doesn’t describe the client API. Both vLLM and SGLang serve an OpenAI-compatible API, as covered in the vLLM docs and the SGLang docs. The examples below read the base URL from an environment variable, so point it at your server. vLLM’s default port is 8000, and the SGLang command above uses 30000.

hello.py
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"),
api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),
)
resp = client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Pro-RL",
messages=[{"role": "user", "content": "Explain sliding-window attention in three sentences."}],
temperature=1.0,
top_p=0.95,
)
print(resp.choices[0].message.content)

temperature=1.0 and top_p=0.95 are the card’s recommended sampling settings.

Agent tool calling

General agents were one of the training domains. The card’s vLLM command sets --tool-call-parser mimo --enable-auto-tool-choice. The SGLang command sets --tool-call-parser mimo but not --enable-auto-tool-choice. The card doesn’t describe what the parser returns. vLLM’s tool calling docs describe parsed calls coming back as structured tool_calls objects, and the example below assumes that behavior.

The card also doesn’t document how MiMo’s chat template handles earlier assistant turns in a multi-turn tool loop. To be safe, the example sends back only the assistant’s content and tool_calls and leaves out any reasoning output. It dispatches calls by function name and caps the number of turns.

agent_tool.py
import json
import os
from openai import OpenAI
client = OpenAI(
base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"),
api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),
)
MODEL = "XiaomiMiMo/MiMo-V2.6-Pro-RL"
MAX_TURNS = 5
def get_order_status(order_id: str) -> dict:
# Replace with a real lookup
return {"order_id": order_id, "status": "shipped", "eta_days": 2}
TOOL_FUNCS = {"get_order_status": get_order_status}
tools = [{
"type": "function",
"function": {
"name": "get_order_status",
"description": "Look up the shipping status of an order.",
"parameters": {
"type": "object",
"properties": {"order_id": {"type": "string"}},
"required": ["order_id"],
},
},
}]
messages = [{"role": "user", "content": "Where is order A-1042 and when will it arrive?"}]
for _ in range(MAX_TURNS):
resp = client.chat.completions.create(
model=MODEL, messages=messages, tools=tools,
temperature=1.0, top_p=0.95,
)
msg = resp.choices[0].message
if not msg.tool_calls:
print(msg.content)
break
# Send back only content and tool calls, not reasoning output
messages.append({
"role": "assistant",
"content": msg.content or "",
"tool_calls": [
{
"id": c.id,
"type": "function",
"function": {"name": c.function.name, "arguments": c.function.arguments},
}
for c in msg.tool_calls
],
})
for call in msg.tool_calls:
func = TOOL_FUNCS.get(call.function.name)
if func is None:
result = {"error": f"unknown tool: {call.function.name}"}
else:
try:
args = json.loads(call.function.arguments)
result = func(**args)
except (json.JSONDecodeError, TypeError) as e:
result = {"error": f"bad arguments: {e}"}
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
else:
print(f"Stopped after {MAX_TURNS} turns without a final answer.")

Long-context repository review

According to the card, the 1M-token window is meant for “long repositories, tool traces, and multi-session agent runs.” Of the 70 layers, 60 use a 128-token sliding window and 10 use global attention over the full context. The context you actually get depends on your deployment. The vLLM command uses --max-model-len auto, and the card doesn’t say what that resolves to on a given setup, so check the server’s reported max length first.

The script below packs a codebase into a single prompt and asks for a review. It stops adding files once a character budget is reached. Characters are only a rough stand-in for tokens, so set MAX_PROMPT_CHARS to fit your server’s limit.

repo_review.py
import os
import pathlib
import sys
from openai import OpenAI
client = OpenAI(
base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"),
api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),
)
MAX_PROMPT_CHARS = int(os.environ.get("MAX_PROMPT_CHARS", "400000"))
repo = pathlib.Path(sys.argv[1] if len(sys.argv) > 1 else ".")
exts = {".py", ".ts", ".js", ".go", ".rs", ".md"}
parts = []
total = 0
skipped = 0
for path in sorted(repo.rglob("*")):
if not (path.is_file() and path.suffix in exts and ".git" not in path.parts):
continue
try:
chunk = f"### FILE: {path.relative_to(repo)}\n{path.read_text()}"
except (UnicodeDecodeError, OSError):
continue
if total + len(chunk) > MAX_PROMPT_CHARS:
skipped += 1
continue
parts.append(chunk)
total += len(chunk)
if skipped:
print(f"Skipped {skipped} files to stay under {MAX_PROMPT_CHARS} characters.", file=sys.stderr)
prompt = (
"Below is a repository. List the three riskiest bugs you can find, "
"with file paths and a suggested fix for each.\n\n" + "\n\n".join(parts)
)
resp = client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Pro-RL",
messages=[{"role": "user", "content": prompt}],
temperature=1.0,
top_p=0.95,
)
print(resp.choices[0].message.content)

Batch processing

The SGLang reference config sets --max-running-requests 128, so that server runs up to 128 requests concurrently. Requests beyond that wait in a queue rather than being refused. To use that capacity from a client, send requests concurrently. If you’re targeting the SGLang server, set MIMO_BASE_URL to http://<host>:30000/v1. The semaphore of 32 below stays under that limit. The card gives no concurrency settings for vLLM, so tune the semaphore for your setup there.

batch_extract.py
import asyncio
import json
import os
from openai import AsyncOpenAI
client = AsyncOpenAI(
base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"),
api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),
)
sem = asyncio.Semaphore(32)
tickets = [
"App crashes when I upload a PNG larger than 10MB on Android 15.",
"Billing page shows EUR but I'm in the US.",
"Login with SSO loops back to the sign-in page.",
]
async def classify(text: str) -> dict:
async with sem:
resp = await client.chat.completions.create(
model="XiaomiMiMo/MiMo-V2.6-Pro-RL",
messages=[{
"role": "user",
"content": (
"Return only JSON with keys 'area' (one of: upload, billing, auth, other) "
f"and 'summary' (max 10 words) for this ticket:\n{text}"
),
}],
temperature=1.0,
top_p=0.95,
)
return {"ticket": text, "result": resp.choices[0].message.content}
async def main():
results = await asyncio.gather(*(classify(t) for t in tickets))
print(json.dumps(results, indent=2))
asyncio.run(main())

Gotchas

  • Hardware requirements are high. The model has 1.02T total parameters. The lightest config on the card uses tensor parallelism of 8, and the SGLang reference setup spans two nodes. The card gives no GPU type, memory figure or weight precision, so work out whether your GPUs can hold it. A single consumer GPU or a laptop won’t run it.
  • Keep --trust-remote-code. Both serving commands on the card include it.
  • Keep the MiMo parsers. Both commands set --reasoning-parser mimo and --tool-call-parser mimo. Leave them in.
  • The recipes are for V2.5. The vLLM image tag (mimov25-cu129) and the cookbook links point to MiMo-V2.5 recipes. Check those pages for any V2.6-specific changes before you rely on them.
  • Multimodal input is undocumented. The card claims image, video and audio support but does not show a request format for them. All examples here are text-only, so check the SGLang/vLLM recipes before you send media.
  • No transformers snippet. The library tag is transformers, but the card documents only SGLang and vLLM deployment.
  • Sampling settings. The card recommends temperature=1.0, top_p=0.95.

When to pick it

Pick MiMo-V2.6-Pro-RL if you need an MIT-licensed model you can self-host for agentic work and you already have multi-GPU infrastructure. Typical workloads are terminal and coding agents, workflow automation, and long tool traces. Its scores on AutomationBench (53.1) and Terminal Bench 2.1 (89.9) are the highest among the models the card compares. The 1M context window is aimed at whole-repository prompts, but check how much of it your deployment actually supports.

If you can’t run it yourself, use a hosted option. The card lists the Xiaomi MiMo API Platform and OpenRouter. MiMo-V2.6-Flash-RL is also worth a look: it comes within a few points of Pro on most agent benchmarks. The card gives no size or hardware requirements for Flash, though, so don’t assume it is lighter to run. Look elsewhere if your main workload is exploit development, where the card shows it well behind GPT-5.6 Sol and Claude Opus 5. The same goes for hard terminal tasks like Terminal Bench 4.0, where it scores 34.9 against Claude Opus 5’s 49.0 and GPT-5.6 Sol’s 39.9. Also look elsewhere if you need documented multimodal input today: the card claims omnimodal support but gives no request examples for it.

Related