Run MiMo-V2.6-Distill-Qwen-9B for Coding and Agent Tasks
Xiaomi's 9B SFT model, fine-tuned from Qwen3.5-9B, scores higher than its base on SWE Pro, Terminal Bench and Toolathlon. How to serve it with SGLang.
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Xiaomi's 9B SFT model, fine-tuned from Qwen3.5-9B, scores higher than its base on SWE Pro, Terminal Bench and Toolathlon. How to serve it with SGLang.
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
MiniMax H3 is an open-weight model that generates video and stereo audio together. Here's how to serve it with SGLang and where the hosted APIs fit in.
MiniMaxAI/MiniMax-H3
NeoHorse-1-4B is a text-only Qwen3.5-4B fine-tune for agent harnesses, tool use and coding. Here is how to serve it with SGLang or vLLM.
TokenRhythm/NeoHorse-1-4B
Nex-AGI's smallest Nex-N2.5 agent model, served with SGLang on 2 H100s, with tool calling and per-request thinking modes.
nex-agi/Nex-N2.5-mini