DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
IFM's open MoE model runs 4B active parameters with a 512K context window. How to serve it, call it, and use it for agents and long documents.
IFM/K2-Horizon-MoVA-36B-A4B
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
MiniCPM5-2B is a 2.5B Llama-architecture model with 131K context and tool calling. Here's how to serve it and what it's good for.
openbmb/MiniCPM5-2B
Qwen3.8-27B is an Apache-2.0 dense 27B vision-language model with adjustable reasoning. How to serve it and call it today.
Qwen/Qwen3.8-27B
China Telecom's 29B MoE activates 4B parameters per token and has a 256K context. Here's how to call it through an OpenAI-compatible API.
XingChen-AGI/Xing4.0-29B-A4B