DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Google DeepMind's largest dense Gemma 4 model handles images, video frames and 256K-token context. Here's how to run it with Transformers.
google/gemma-4-31B-it
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
A 27B vision decision model that returns a calibrated probability for every option in one forward pass. Here's how to serve it and use it.
autotrust/JEV-27B-VL
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
MiniMax H3 is an open-weight model that generates video and stereo audio together. Here's how to serve it with SGLang and where the hosted APIs fit in.
MiniMaxAI/MiniMax-H3
Qwen's driving model adds a BEV perception head and a flow-matching planner to Qwen3.5-4B. Install it and sample ego trajectories from demo scenes.
Qwen/Qwen-Drive-1.0-4B
TeleOCR is a 1.2B Apache-2.0 model that turns text, tables, formulas and layouts into structured output, including from photos of warped pages.
XingChen-AGI/TeleOCR
Clef answers typed questions about any input (text, JSON, images) with a probability per option in one forward pass. Here's how to use it for routing, triage and classification.
Cloudflare/clef
Google's new 740M multimodal embedding model runs on a CPU. Here's how to build a working semantic search over your own documents, and the prefixes that make it accurate.
google/embeddinggemma-2