DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
DeepSeek-V4.1-Flash is DeepSeek-AI’s new open-weight multimodal Mixture-of-Experts model. It takes images and text as input, generates text, and supports contexts of up to one million tokens. Its main feature is a much smaller KV cache. Compressed Sparse Attention 2 (CSA2) and FP4 main KV caching bring the global KV cache down to 890 bytes per token. That is roughly 1/4 of DeepSeek-V4-Flash and, according to the card, about 437x smaller than DeepSeek-V1. Separately, a Causal Encoder-Decoder (CED) layout means the model activates only 8B parameters per token during prefill, so it targets agent workloads that read far more than they write. The weights are MIT licensed.
Key specs
| Spec | Value |
|---|---|
| Backbone parameters | 552B (MoE) |
| Activated parameters | 8B prefill / 16B decode |
| Layers | 40 (20-layer causal encoder + 20-layer decoder) |
| Experts per MoE layer | 1 shared + 384 routed, 6 routed active per token |
| Engram conditional memory | 196B parameters, sparsely accessed |
| Context length | up to 1M tokens |
| Global KV cache | 890 bytes/token (FP4 E2M1 main KV) |
| Persistent KV footprint | ~1/8 of DeepSeek-V4-Flash (via SWA Bounded Replay) |
| Pre-training data | 45T multimodal tokens |
| Vision | DeepSeek-ViT (trained from scratch, 2D-RoPE) + 2-layer MLP projector |
| Reasoning effort | integer 1–100, continuously controllable |
| License | MIT |
The CED design projects the decoder’s global KV cache from the final encoder hidden states instead of computing it in every decoder layer. That design is the reason prefill costs 8B active parameters while decode costs 16B. The model also includes DSpark speculative decoding and a Hierarchical Sparse Indexer, which keeps the cost of deeper indexing layers from growing with context length.
On instruct benchmarks at max reasoning effort, the card reports these scores:
- Terminal-Bench 2.1: 90.6, the top score in the comparison table (Opus-5.0 scores 89.1).
- DeepSWE v1.1: 74.2 resolved, next to Opus-5.0 at 74.0 and GPT-5.6 Sol at 73.0.
- AutomationBench: 54.8. Agent’s Last Exam: 31.8. CyberGym: 88.1. HLE with tools: 63.9. Codeforces rating: 3471.
These headline agent numbers come from specific harnesses: the card evaluates DeepSWE v1.1 with mini-SWE and Terminal-Bench 2.1 with DeepSeek Harness (DSH) Minimal. The card doesn’t say which harness the other models used. In Claude Code or Codex, V4.1-Flash scores 65.6–69.8 on DeepSWE, below the Opus-5.0 and GPT-5.6 Sol figures above (see the scaffold table below).
The model is not ahead everywhere:
- Terminal-Bench 3.0 and 4.0: 30.0 and 31.2, against 43.3 and 51.8 for Opus-5.0.
- ProgramBench: 20.3, against 37.0 for Opus-5.0.
- SEC-Bench Pro: 62.8, against 74.3 for GPT-5.6 Sol.
- ExploitGym: 15.3, against 33.7 for GPT-5.6 Sol and 22.1 for Opus-5.0.
- HLE without tools: 36.8.
Install
The card does not give a pip install line or a transformers loading snippet. The Hub metadata says library_name: transformers, but local inference goes through the repo’s own tooling:
- Weights and local inference: the
inferencefolder documents weight conversion and how to run the model locally. - Prompt formatting: this release has no Jinja chat template. The
encodingfolder has a self-contained reference implementation (encoding.py) with test cases. - Production prompt handling: deepseek-recipe is a set of Rust libraries with Python bindings. They convert Messages, Chat Completions and Responses API requests into DeepSeek’s Conversation format, encode them into V4/V4.1 prompts or token IDs, and parse output back into full or streamed responses.
We are not reproducing the conversion or loading code here because the model card itself doesn’t show it. Follow inference/README.md in the repo for the exact commands.
Run it
The card gives recommended sampling settings. Use them whatever serving stack you end up with:
| Parameter | Value |
|---|---|
temperature |
1.0 |
top_p |
0.95 or 1.0 |
context_window |
1M tokens |
max_tokens |
≥ 256K |
All of the card’s instruct results use reasoning_effort=100 with temperature=1.0, top_p=0.95. If you want to reproduce the published numbers, start from those values. The card doesn’t explain why it recommends a max_tokens of 256K or more. Our guess is that the model needs room for long reasoning at high effort settings. Check what output limit your client sets by default and raise it if needed.
Coding agents in an existing harness
The card gives more detail on this use case than on any other. It tests the same model across eight agent scaffolds with N=8 samples on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1, using a 1M-token context and max_steps=500:
| Scaffold | DeepSWE v1.1 | Terminal-Bench 2.1 |
|---|---|---|
| Claude Code | 69.8 | 88.0 |
| Codex | 65.6 | 84.1 |
| OpenCode | 65.5 | 85.0 |
| Pi | 66.2 | 86.1 |
| mini-SWE | 74.2 | 90.3 |
| DSH Minimal | 72.6 | 90.6 |
| DSH Standard | 70.5 | 85.8 |
| DSH PTC | 67.6 | 85.8 |
Scores change by up to 8.7 points on DeepSWE depending on the harness, so the scaffold matters. mini-SWE and DSH Minimal scored highest on both benchmarks in these runs. In tools many teams already use, the scores are lower: Claude Code reaches 69.8 on DeepSWE and 88.0 on Terminal-Bench 2.1, and OpenCode reaches 65.5 and 85.0.
The evaluation folder walks through reproducing the DeepSWE v1.1 results with both dsh-minimal and the official mini-swe-agent. It also includes the patch for integrating dsh-minimal with Pier. Running that evaluation is the best starting point if you want to check the model on your own infrastructure before switching an agent over.
Long-document and repository-scale input
The card says CED improves cost efficiency for input-heavy agentic workloads, meaning prompts that are mostly input: whole repositories, long logs, document sets. Prefill activates 8B parameters per token, and the global KV cache costs 890 bytes per token. So holding very long contexts takes much less KV memory than with previous DeepSeek generations. The card trained sparse attention at 64K and extended context to 1M tokens partway through pre-training, at the 34T-token mark.
Keep expectations calibrated. On the long-context base benchmark, LongBench-V2, V4.1-Flash-Base scores 45.2, which is close to V4-Flash-Base at 44.7 and below V4-Pro-Base at 51.5. The improvement is in what long context costs. The card doesn’t show a jump in long-context accuracy.
Document and chart understanding
Images and text are trained jointly from the start of pre-training, not added afterward. The card reports these multimodal scores for the base model (4-shot, except RefCOCO at 0-shot):
- DocVQA (LLM-Judge): 95.6
- CVBench: 77.9
- MMMU-Pro: 56.5
- RefCOCO-avg (Acc@0.5): 86.0
The card gives no instruct-model numbers for these benchmarks, so they don’t directly tell you what the deployed instruct model will score.
For the instruct model, the card reports visual-agent results with tools: 78.9 on Chartography, 89.6 on BabyVision and 49.0 Pass@5 on ZeroBench-main. These are evaluated in the Claude Code harness with a 512K context. Four models have scores on all three benchmarks: Opus-5.0, GPT-5.6 Sol, K3 and V4.1-Flash. On BabyVision, V4.1-Flash is second, behind Opus-5.0 (94.1) and ahead of GPT-5.6 Sol (88.9) and K3 (85.7). On Chartography and ZeroBench-main it is third of four. Opus-5.0 (84.0) and GPT-5.6 Sol (79.9) beat it on Chartography, and GPT-5.6 Sol (53.0) and Opus-5.0 (52.0) beat it on ZeroBench-main. The encoding.py reference covers interleaved image content in prompts, so use it to check your message format.
Gotchas
- No chat template.
tokenizer.apply_chat_templatestyle workflows won’t work out of the box. Format prompts withencoding/encoding.pyor deepseek-recipe. That includes tool calls, thinking mode, numeric reasoning effort and mid-conversation system messages. - Weight conversion. The card points to
inference/README.mdfor instructions on weight conversion and running inference locally. Read it before you plan hardware. - Size. The card lists 552B backbone parameters and a 196B-parameter Engram conditional memory. It doesn’t say whether Engram is counted within the 552B. Only 8B–16B are active per token, but all the weights still have to be stored. The card does not list minimum hardware.
- Reasoning effort trades cost for accuracy. Published scores use effort 100. Lower settings cost less, but the card does not report accuracy at lower settings, so measure on your own tasks.
- Huge output budgets. The recommended
max_tokensis 256K or more. Plan latency and cost around that, or test how much quality you lose with a lower cap. - Harness sensitivity. Agent results vary by several points across scaffolds (see the table above), so a quick test in one harness may not reflect what the model can do in another.
When to pick it
Pick it when your workload is agentic and input-heavy, such as coding agents or terminal automation, and you want open MIT weights. It has the top score in its comparison table on Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, AutomationBench and Agent’s Last Exam, though the DeepSWE and Terminal-Bench 2.1 figures come from its best-scoring harnesses. It is also a good fit when KV cache memory limits how many long sessions you can serve at once. A per-token cache about 4x smaller than V4-Flash’s can matter for serving cost.
Look elsewhere in these cases:
- Your tasks look like the harder agent suites. On Terminal-Bench 3.0/4.0 and ProgramBench it trails Opus-5.0 by a wide margin.
- Your work is security-focused beyond CyberGym. On SEC-Bench Pro (62.8 vs 74.3) and ExploitGym (15.3 vs 33.7) it trails GPT-5.6 Sol by a wide margin.
- You need the strongest visual-agent results. Opus-5.0 beats it on Chartography, BabyVision and ZeroBench-main, and GPT-5.6 Sol beats it on Chartography and ZeroBench-main.
- You need the best knowledge recall. V4-Pro-Base scores well above it on SimpleQA-Verified (55.2 vs 42.3).
- Your prompts are multilingual math. V4.1-Flash-Base is the lowest of the three base models on MGSM at 80.2.
- You need something you can
pip installand load in a few lines. The custom prompt encoding and weight-conversion step make this a model for teams ready to run their own serving stack.
Related
GLM-5.3-Flash: What You Can Run With Z.ai's 18B-Active Multimodal Model
Z.ai's first natively multimodal GLM-5 model: 320B total / 18B active parameters, hybrid sparse-linear attention, MIT license. What the card covers.
zai-org/GLM-5.3-Flash
Run Kimi K3: Moonshot's 2.8T Open MoE for Agents and Coding
Kimi K3 is a 2.8T-parameter open-weight multimodal MoE with a 1M-token context. What the model card says and how to call it.
moonshotai/Kimi-K3
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL