Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 7 min read

DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work

DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

DeepSeek-V4.1-Flash is DeepSeek-AI’s new open-weight multimodal Mixture-of-Experts model. It takes images and text as input, generates text, and supports contexts of up to one million tokens. Its main feature is a much smaller KV cache. Compressed Sparse Attention 2 (CSA2) and FP4 main KV caching bring the global KV cache down to 890 bytes per token. That is roughly 1/4 of DeepSeek-V4-Flash and, according to the card, about 437x smaller than DeepSeek-V1. Separately, a Causal Encoder-Decoder (CED) layout means the model activates only 8B parameters per token during prefill, so it targets agent workloads that read far more than they write. The weights are MIT licensed.

Key specs

Spec Value
Backbone parameters 552B (MoE)
Activated parameters 8B prefill / 16B decode
Layers 40 (20-layer causal encoder + 20-layer decoder)
Experts per MoE layer 1 shared + 384 routed, 6 routed active per token
Engram conditional memory 196B parameters, sparsely accessed
Context length up to 1M tokens
Global KV cache 890 bytes/token (FP4 E2M1 main KV)
Persistent KV footprint ~1/8 of DeepSeek-V4-Flash (via SWA Bounded Replay)
Pre-training data 45T multimodal tokens
Vision DeepSeek-ViT (trained from scratch, 2D-RoPE) + 2-layer MLP projector
Reasoning effort integer 1–100, continuously controllable
License MIT

The CED design projects the decoder’s global KV cache from the final encoder hidden states instead of computing it in every decoder layer. That design is the reason prefill costs 8B active parameters while decode costs 16B. The model also includes DSpark speculative decoding and a Hierarchical Sparse Indexer, which keeps the cost of deeper indexing layers from growing with context length.

On instruct benchmarks at max reasoning effort, the card reports these scores:

  • Terminal-Bench 2.1: 90.6, the top score in the comparison table (Opus-5.0 scores 89.1).
  • DeepSWE v1.1: 74.2 resolved, next to Opus-5.0 at 74.0 and GPT-5.6 Sol at 73.0.
  • AutomationBench: 54.8. Agent’s Last Exam: 31.8. CyberGym: 88.1. HLE with tools: 63.9. Codeforces rating: 3471.

These headline agent numbers come from specific harnesses: the card evaluates DeepSWE v1.1 with mini-SWE and Terminal-Bench 2.1 with DeepSeek Harness (DSH) Minimal. The card doesn’t say which harness the other models used. In Claude Code or Codex, V4.1-Flash scores 65.6–69.8 on DeepSWE, below the Opus-5.0 and GPT-5.6 Sol figures above (see the scaffold table below).

The model is not ahead everywhere:

  • Terminal-Bench 3.0 and 4.0: 30.0 and 31.2, against 43.3 and 51.8 for Opus-5.0.
  • ProgramBench: 20.3, against 37.0 for Opus-5.0.
  • SEC-Bench Pro: 62.8, against 74.3 for GPT-5.6 Sol.
  • ExploitGym: 15.3, against 33.7 for GPT-5.6 Sol and 22.1 for Opus-5.0.
  • HLE without tools: 36.8.

Install

The card does not give a pip install line or a transformers loading snippet. The Hub metadata says library_name: transformers, but local inference goes through the repo’s own tooling:

  • Weights and local inference: the inference folder documents weight conversion and how to run the model locally.
  • Prompt formatting: this release has no Jinja chat template. The encoding folder has a self-contained reference implementation (encoding.py) with test cases.
  • Production prompt handling: deepseek-recipe is a set of Rust libraries with Python bindings. They convert Messages, Chat Completions and Responses API requests into DeepSeek’s Conversation format, encode them into V4/V4.1 prompts or token IDs, and parse output back into full or streamed responses.

We are not reproducing the conversion or loading code here because the model card itself doesn’t show it. Follow inference/README.md in the repo for the exact commands.

Run it

The card gives recommended sampling settings. Use them whatever serving stack you end up with:

Parameter Value
temperature 1.0
top_p 0.95 or 1.0
context_window 1M tokens
max_tokens ≥ 256K

All of the card’s instruct results use reasoning_effort=100 with temperature=1.0, top_p=0.95. If you want to reproduce the published numbers, start from those values. The card doesn’t explain why it recommends a max_tokens of 256K or more. Our guess is that the model needs room for long reasoning at high effort settings. Check what output limit your client sets by default and raise it if needed.

Coding agents in an existing harness

The card gives more detail on this use case than on any other. It tests the same model across eight agent scaffolds with N=8 samples on DeepSWE v1.1 and N=3 on Terminal-Bench 2.1, using a 1M-token context and max_steps=500:

Scaffold DeepSWE v1.1 Terminal-Bench 2.1
Claude Code 69.8 88.0
Codex 65.6 84.1
OpenCode 65.5 85.0
Pi 66.2 86.1
mini-SWE 74.2 90.3
DSH Minimal 72.6 90.6
DSH Standard 70.5 85.8
DSH PTC 67.6 85.8

Scores change by up to 8.7 points on DeepSWE depending on the harness, so the scaffold matters. mini-SWE and DSH Minimal scored highest on both benchmarks in these runs. In tools many teams already use, the scores are lower: Claude Code reaches 69.8 on DeepSWE and 88.0 on Terminal-Bench 2.1, and OpenCode reaches 65.5 and 85.0.

The evaluation folder walks through reproducing the DeepSWE v1.1 results with both dsh-minimal and the official mini-swe-agent. It also includes the patch for integrating dsh-minimal with Pier. Running that evaluation is the best starting point if you want to check the model on your own infrastructure before switching an agent over.

Long-document and repository-scale input

The card says CED improves cost efficiency for input-heavy agentic workloads, meaning prompts that are mostly input: whole repositories, long logs, document sets. Prefill activates 8B parameters per token, and the global KV cache costs 890 bytes per token. So holding very long contexts takes much less KV memory than with previous DeepSeek generations. The card trained sparse attention at 64K and extended context to 1M tokens partway through pre-training, at the 34T-token mark.

Keep expectations calibrated. On the long-context base benchmark, LongBench-V2, V4.1-Flash-Base scores 45.2, which is close to V4-Flash-Base at 44.7 and below V4-Pro-Base at 51.5. The improvement is in what long context costs. The card doesn’t show a jump in long-context accuracy.

Document and chart understanding

Images and text are trained jointly from the start of pre-training, not added afterward. The card reports these multimodal scores for the base model (4-shot, except RefCOCO at 0-shot):

  • DocVQA (LLM-Judge): 95.6
  • CVBench: 77.9
  • MMMU-Pro: 56.5
  • RefCOCO-avg (Acc@0.5): 86.0

The card gives no instruct-model numbers for these benchmarks, so they don’t directly tell you what the deployed instruct model will score.

For the instruct model, the card reports visual-agent results with tools: 78.9 on Chartography, 89.6 on BabyVision and 49.0 Pass@5 on ZeroBench-main. These are evaluated in the Claude Code harness with a 512K context. Four models have scores on all three benchmarks: Opus-5.0, GPT-5.6 Sol, K3 and V4.1-Flash. On BabyVision, V4.1-Flash is second, behind Opus-5.0 (94.1) and ahead of GPT-5.6 Sol (88.9) and K3 (85.7). On Chartography and ZeroBench-main it is third of four. Opus-5.0 (84.0) and GPT-5.6 Sol (79.9) beat it on Chartography, and GPT-5.6 Sol (53.0) and Opus-5.0 (52.0) beat it on ZeroBench-main. The encoding.py reference covers interleaved image content in prompts, so use it to check your message format.

Gotchas

  • No chat template. tokenizer.apply_chat_template style workflows won’t work out of the box. Format prompts with encoding/encoding.py or deepseek-recipe. That includes tool calls, thinking mode, numeric reasoning effort and mid-conversation system messages.
  • Weight conversion. The card points to inference/README.md for instructions on weight conversion and running inference locally. Read it before you plan hardware.
  • Size. The card lists 552B backbone parameters and a 196B-parameter Engram conditional memory. It doesn’t say whether Engram is counted within the 552B. Only 8B–16B are active per token, but all the weights still have to be stored. The card does not list minimum hardware.
  • Reasoning effort trades cost for accuracy. Published scores use effort 100. Lower settings cost less, but the card does not report accuracy at lower settings, so measure on your own tasks.
  • Huge output budgets. The recommended max_tokens is 256K or more. Plan latency and cost around that, or test how much quality you lose with a lower cap.
  • Harness sensitivity. Agent results vary by several points across scaffolds (see the table above), so a quick test in one harness may not reflect what the model can do in another.

When to pick it

Pick it when your workload is agentic and input-heavy, such as coding agents or terminal automation, and you want open MIT weights. It has the top score in its comparison table on Terminal-Bench 2.1, DeepSWE v1.1, CyberGym, AutomationBench and Agent’s Last Exam, though the DeepSWE and Terminal-Bench 2.1 figures come from its best-scoring harnesses. It is also a good fit when KV cache memory limits how many long sessions you can serve at once. A per-token cache about 4x smaller than V4-Flash’s can matter for serving cost.

Look elsewhere in these cases:

  • Your tasks look like the harder agent suites. On Terminal-Bench 3.0/4.0 and ProgramBench it trails Opus-5.0 by a wide margin.
  • Your work is security-focused beyond CyberGym. On SEC-Bench Pro (62.8 vs 74.3) and ExploitGym (15.3 vs 33.7) it trails GPT-5.6 Sol by a wide margin.
  • You need the strongest visual-agent results. Opus-5.0 beats it on Chartography, BabyVision and ZeroBench-main, and GPT-5.6 Sol beats it on Chartography and ZeroBench-main.
  • You need the best knowledge recall. V4-Pro-Base scores well above it on SimpleQA-Verified (55.2 vs 42.3).
  • Your prompts are multilingual math. V4.1-Flash-Base is the lowest of the three base models on MGSM at 80.2.
  • You need something you can pip install and load in a few lines. The custom prompt encoding and weight-conversion step make this a model for teams ready to run their own serving stack.

Related