Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
MiMo-V2.6-Pro-RL is the flagship checkpoint of Xiaomi’s MiMo-V2.6 series. It is a sparse Mixture-of-Experts model with 1.02T total parameters, 42B of them active per token. It takes text, image, video and audio, and its context window is 1M tokens. Most of the card is about how the model was trained. Xiaomi ran one mixed reinforcement-learning run across coding, general agents, visual tasks and cybersecurity, rather than one run per domain. A “groupwise agentic grader” then ranks passing solutions against each other instead of scoring them only pass/fail. The model is released under MIT. On several agent benchmarks it scores close to the closed models the card compares it with. On others it is well behind, including Terminal Bench 4.0 and the exploit benchmarks.
Key specs
| MiMo-V2.6-Pro-RL | |
|---|---|
| Architecture | Sparse MoE, 1.02T total / 42B activated |
| Experts | 384 routed, 8 activated, no shared experts |
| Layers | 70 (60 sliding-window, 10 global attention) |
| Sliding window | 128 tokens |
| Context | 1M tokens |
| Modalities | Text, image, video, audio |
| Vision encoder | 681M-param MiMo ViT |
| Audio | 308M AudioTokenizer + 127M audio patch encoder |
| Speculative decoder | 5-layer MTP drafter, predicts 7 tokens per pass |
| License | MIT |
These are selected benchmark results from the card. The “Pro” column is this model, and the others are the comparison models the card lists:
| Benchmark | V2.6 Pro | V2.6 Flash | Claude Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|
| DeepSWE v1.1 | 71.9 | 67.9 | 74.0 | 73.0 |
| AutomationBench v1.0.6 | 53.1 | 52.3 | 50.3 | 45.8 |
| Toolathlon-Verified | 76.9 | 73.6 | 80.6 | 74.9 |
| Terminal Bench 2.1 | 89.9 | 87.6 | 89.1 | 88.8 |
| Terminal Bench 4.0 | 34.9 | 28.8 | 49.0 | 39.9 |
| OSWorld-Verified | 82.0 | 80.8 | 83.4 | 83.0 |
| ExploitBench | 47.9 | 25.3 | 70.0 | 78.5 |
Among the models in this table, it has the top score on AutomationBench and Terminal Bench 2.1. It trails Claude Opus 5 on DeepSWE, Toolathlon, Terminal Bench 4.0 and OSWorld-Verified. It also trails GPT-5.6 Sol on DeepSWE (71.9 vs 73.0), Terminal Bench 4.0 (34.9 vs 39.9) and OSWorld-Verified (82.0 vs 83.0). It is well behind both closed models on ExploitBench. Compared with the previous MiMo-V2.5 Pro, the gains are large: DeepSWE went from 19.0 to 71.9 and AutomationBench from 16.0 to 53.1.
Install
The card gives two serving paths and names a prebuilt Docker image for each. The vLLM image and both linked recipes are MiMo-V2.5 artifacts: the vLLM tag is mimov25-cu129, and the cookbook and recipe pages are for V2.5. The SGLang image is the generic lmsysorg/sglang:latest. Check the recipe pages for any V2.6-specific changes before you rely on them.
# vLLM image (from the vLLM MiMo-V2.5 recipe)docker pull vllm/vllm-openai:mimov25-cu129
# SGLang image (named alongside the SGLang MiMo cookbook)docker pull lmsysorg/sglang:latestThe card doesn’t include a docker run command, so it doesn’t cover GPU flags, port mapping or cache mounts. Follow the vLLM MiMo-V2.5 recipe or the SGLang MiMo cookbook to start the container. Run the serve commands below inside that container, or on a host where vLLM or SGLang is already installed.
To query the server from Python, install the OpenAI client:
pip install openaiRun it
The card’s vLLM command is the shorter of the two. It uses tensor parallelism of 8:
vllm serve XiaomiMiMo/MiMo-V2.6-Pro-RL \ --tensor-parallel-size 8 \ --trust-remote-code \ --gpu-memory-utilization 0.95 \ --max-model-len auto \ --reasoning-parser mimo \ --tool-call-parser mimo \ --enable-auto-tool-choice \ --generation-config vllmThe card gives no GPU type, memory figure or weight precision. Check that a 1.02T-parameter model fits on your 8 GPUs before you start.
“For best performance,” the card points to the SGLang MiMo cookbook. Xiaomi’s reference SGLang setup runs across 2 nodes with TP=16, DP=2, expert parallelism, DeepEP all-to-all and EAGLE speculative decoding. The full command is in the model card and the SGLang MiMo cookbook. It listens on port 30000.
The model card doesn’t describe the client API. Both vLLM and SGLang serve an OpenAI-compatible API, as covered in the vLLM docs and the SGLang docs. The examples below read the base URL from an environment variable, so point it at your server. vLLM’s default port is 8000, and the SGLang command above uses 30000.
import osfrom openai import OpenAI
client = OpenAI( base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"), api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),)
resp = client.chat.completions.create( model="XiaomiMiMo/MiMo-V2.6-Pro-RL", messages=[{"role": "user", "content": "Explain sliding-window attention in three sentences."}], temperature=1.0, top_p=0.95,)print(resp.choices[0].message.content)temperature=1.0 and top_p=0.95 are the card’s recommended sampling settings.
Agent tool calling
General agents were one of the training domains. The card’s vLLM command sets --tool-call-parser mimo --enable-auto-tool-choice. The SGLang command sets --tool-call-parser mimo but not --enable-auto-tool-choice. The card doesn’t describe what the parser returns. vLLM’s tool calling docs describe parsed calls coming back as structured tool_calls objects, and the example below assumes that behavior.
The card also doesn’t document how MiMo’s chat template handles earlier assistant turns in a multi-turn tool loop. To be safe, the example sends back only the assistant’s content and tool_calls and leaves out any reasoning output. It dispatches calls by function name and caps the number of turns.
import jsonimport osfrom openai import OpenAI
client = OpenAI( base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"), api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),)MODEL = "XiaomiMiMo/MiMo-V2.6-Pro-RL"MAX_TURNS = 5
def get_order_status(order_id: str) -> dict: # Replace with a real lookup return {"order_id": order_id, "status": "shipped", "eta_days": 2}
TOOL_FUNCS = {"get_order_status": get_order_status}
tools = [{ "type": "function", "function": { "name": "get_order_status", "description": "Look up the shipping status of an order.", "parameters": { "type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"], }, },}]
messages = [{"role": "user", "content": "Where is order A-1042 and when will it arrive?"}]
for _ in range(MAX_TURNS): resp = client.chat.completions.create( model=MODEL, messages=messages, tools=tools, temperature=1.0, top_p=0.95, ) msg = resp.choices[0].message if not msg.tool_calls: print(msg.content) break
# Send back only content and tool calls, not reasoning output messages.append({ "role": "assistant", "content": msg.content or "", "tool_calls": [ { "id": c.id, "type": "function", "function": {"name": c.function.name, "arguments": c.function.arguments}, } for c in msg.tool_calls ], })
for call in msg.tool_calls: func = TOOL_FUNCS.get(call.function.name) if func is None: result = {"error": f"unknown tool: {call.function.name}"} else: try: args = json.loads(call.function.arguments) result = func(**args) except (json.JSONDecodeError, TypeError) as e: result = {"error": f"bad arguments: {e}"} messages.append({ "role": "tool", "tool_call_id": call.id, "content": json.dumps(result), })else: print(f"Stopped after {MAX_TURNS} turns without a final answer.")Long-context repository review
According to the card, the 1M-token window is meant for “long repositories, tool traces, and multi-session agent runs.” Of the 70 layers, 60 use a 128-token sliding window and 10 use global attention over the full context. The context you actually get depends on your deployment. The vLLM command uses --max-model-len auto, and the card doesn’t say what that resolves to on a given setup, so check the server’s reported max length first.
The script below packs a codebase into a single prompt and asks for a review. It stops adding files once a character budget is reached. Characters are only a rough stand-in for tokens, so set MAX_PROMPT_CHARS to fit your server’s limit.
import osimport pathlibimport sysfrom openai import OpenAI
client = OpenAI( base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"), api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),)
MAX_PROMPT_CHARS = int(os.environ.get("MAX_PROMPT_CHARS", "400000"))
repo = pathlib.Path(sys.argv[1] if len(sys.argv) > 1 else ".")exts = {".py", ".ts", ".js", ".go", ".rs", ".md"}parts = []total = 0skipped = 0for path in sorted(repo.rglob("*")): if not (path.is_file() and path.suffix in exts and ".git" not in path.parts): continue try: chunk = f"### FILE: {path.relative_to(repo)}\n{path.read_text()}" except (UnicodeDecodeError, OSError): continue if total + len(chunk) > MAX_PROMPT_CHARS: skipped += 1 continue parts.append(chunk) total += len(chunk)
if skipped: print(f"Skipped {skipped} files to stay under {MAX_PROMPT_CHARS} characters.", file=sys.stderr)
prompt = ( "Below is a repository. List the three riskiest bugs you can find, " "with file paths and a suggested fix for each.\n\n" + "\n\n".join(parts))
resp = client.chat.completions.create( model="XiaomiMiMo/MiMo-V2.6-Pro-RL", messages=[{"role": "user", "content": prompt}], temperature=1.0, top_p=0.95,)print(resp.choices[0].message.content)Batch processing
The SGLang reference config sets --max-running-requests 128, so that server runs up to 128 requests concurrently. Requests beyond that wait in a queue rather than being refused. To use that capacity from a client, send requests concurrently. If you’re targeting the SGLang server, set MIMO_BASE_URL to http://<host>:30000/v1. The semaphore of 32 below stays under that limit. The card gives no concurrency settings for vLLM, so tune the semaphore for your setup there.
import asyncioimport jsonimport osfrom openai import AsyncOpenAI
client = AsyncOpenAI( base_url=os.environ.get("MIMO_BASE_URL", "http://localhost:8000/v1"), api_key=os.environ.get("MIMO_API_KEY", "EMPTY"),)sem = asyncio.Semaphore(32)
tickets = [ "App crashes when I upload a PNG larger than 10MB on Android 15.", "Billing page shows EUR but I'm in the US.", "Login with SSO loops back to the sign-in page.",]
async def classify(text: str) -> dict: async with sem: resp = await client.chat.completions.create( model="XiaomiMiMo/MiMo-V2.6-Pro-RL", messages=[{ "role": "user", "content": ( "Return only JSON with keys 'area' (one of: upload, billing, auth, other) " f"and 'summary' (max 10 words) for this ticket:\n{text}" ), }], temperature=1.0, top_p=0.95, ) return {"ticket": text, "result": resp.choices[0].message.content}
async def main(): results = await asyncio.gather(*(classify(t) for t in tickets)) print(json.dumps(results, indent=2))
asyncio.run(main())Gotchas
- Hardware requirements are high. The model has 1.02T total parameters. The lightest config on the card uses tensor parallelism of 8, and the SGLang reference setup spans two nodes. The card gives no GPU type, memory figure or weight precision, so work out whether your GPUs can hold it. A single consumer GPU or a laptop won’t run it.
- Keep
--trust-remote-code. Both serving commands on the card include it. - Keep the MiMo parsers. Both commands set
--reasoning-parser mimoand--tool-call-parser mimo. Leave them in. - The recipes are for V2.5. The vLLM image tag (
mimov25-cu129) and the cookbook links point to MiMo-V2.5 recipes. Check those pages for any V2.6-specific changes before you rely on them. - Multimodal input is undocumented. The card claims image, video and audio support but does not show a request format for them. All examples here are text-only, so check the SGLang/vLLM recipes before you send media.
- No transformers snippet. The library tag is
transformers, but the card documents only SGLang and vLLM deployment. - Sampling settings. The card recommends
temperature=1.0,top_p=0.95.
When to pick it
Pick MiMo-V2.6-Pro-RL if you need an MIT-licensed model you can self-host for agentic work and you already have multi-GPU infrastructure. Typical workloads are terminal and coding agents, workflow automation, and long tool traces. Its scores on AutomationBench (53.1) and Terminal Bench 2.1 (89.9) are the highest among the models the card compares. The 1M context window is aimed at whole-repository prompts, but check how much of it your deployment actually supports.
If you can’t run it yourself, use a hosted option. The card lists the Xiaomi MiMo API Platform and OpenRouter. MiMo-V2.6-Flash-RL is also worth a look: it comes within a few points of Pro on most agent benchmarks. The card gives no size or hardware requirements for Flash, though, so don’t assume it is lighter to run. Look elsewhere if your main workload is exploit development, where the card shows it well behind GPT-5.6 Sol and Claude Opus 5. The same goes for hard terminal tasks like Terminal Bench 4.0, where it scores 34.9 against Claude Opus 5’s 49.0 and GPT-5.6 Sol’s 39.9. Also look elsewhere if you need documented multimodal input today: the card claims omnimodal support but gives no request examples for it.
Related
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
GLM-5.3-Flash: What You Can Run With Z.ai's 18B-Active Multimodal Model
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
Run Kimi K3: Moonshot's 2.8T Open MoE for Agents and Coding
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3