DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Edge0-35B-A3B streams experts from SSD so a 4-bit Qwen3.6-35B-A3B derivative decodes at about 15 tok/s in under 3 GiB of active memory.
Edge0/Edge0-35B-A3B-preview
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
Qwen's 125B MoE with 6B active params handles text, images and video. What it's good at, how to call it, and where it falls short.
Qwen/Qwen3.8-Flash-Next
China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.
XingChen-AGI/Xing4.0-29B-A4B
Kolibri 1 is a German–English reasoning model with only 3.5B active parameters. How to serve it, control its thinking effort, and wire up tool calling.
Aleph-Alpha/Kolibri-1