DeepSeek posted the DeepSeek-V4.1-Flash technical report to arXiv on 17 September 2026. The title is "DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression". The 51-page paper describes the model that DeepSeek released on its API and on Hugging Face on 10 September.
The report makes one central engineering claim. It says the model stores 890 bytes of global KV cache per token, about a quarter of DeepSeek-V4-Flash. The rest of the paper argues that the smaller cache costs little in quality. We read it for what DeepSeek measured, what it compared against, and what it left out.
890 bytes per token and a split between prefill and decode
The paper claims four new pieces. Each one targets memory or compute during serving.
The first piece is the Causal Encoder-Decoder (CED). The model has 40 layers. The bottom 20 layers act as an encoder. The top 20 layers take their KV entries from the encoder's output. The paper says the design is inspired by YOCO, an earlier method that shares KV cache between layer halves. DeepSeek says CED cuts nearly half of prefill computation. As a result, the model activates 8B parameters per token during prefill and 16B during decode.
The second piece is Compressed Sparse Attention 2 (CSA2). Each attention layer runs in one of three fixed modes, called Full, Reindex and Reuse. Reuse layers borrow the KV cache and sparse-attention indices from an earlier layer. A Hierarchical Sparse Indexer limits later layers to a candidate pool. The pool holds up to 2,048 blocks of 8 positions, or 16,384 positions. V4.1-Flash uses CSA2 throughout. DeepSeek-V4 used a hybrid of CSA and Heavily Compressed Attention (HCA).
The third piece is FP4 KV caching. The main KV cache uses the E2M1 format with one E4M3 scale per 16 channels. DeepSeek trained the model with FP4 KV caches and says this caused "only marginal performance degradation."
The fourth piece is a serving change called SWA Bounded Replay. Every layer also uses sliding-window attention (SWA) with a window of 128 tokens. V4 stored those SWA caches on SSD for reuse. V4.1 drops them from the persistent cache. When a request needs them again, the server replays only the last 128 tokens. Exact reconstruction would need the last 40 times 128 tokens. DeepSeek says the approximation causes "only negligible performance degradation."
The paper states the combined result in its abstract. Global KV cache, which always sits in accelerator memory (HBM), falls to 890 bytes per token. Persistent KV cache, which sits on SSD or in host memory, falls to about one eighth of V4-Flash.

Figure 1(b) from the report. Global KV cache per token falls from 389,120 bytes in DeepSeek-V1 to 3,514 bytes in V4-Flash and 890 bytes in V4.1-Flash. Source: DeepSeek-V4.1-Flash technical report, arXiv:2609.19969.
The chart puts V4.1-Flash at a 3.9-fold cut from V4-Flash. The paper calls it about four-fold and 437-fold against DeepSeek-V1.
Our arithmetic shows what that means per request. A full one-million-token context needs about 0.89 GB of global KV cache on V4.1-Flash. The same context on V4-Flash needs about 3.5 GB. These figures cover global KV only. They leave out the small SWA cache and any serving overhead.
The model is large in total. It has 552B backbone parameters and 196B more in Engram, a conditional memory module. Each MoE layer has 1 shared expert and 384 routed experts, with 6 routed experts active per token. The context window is one million tokens. The paper also adds Single-Pass mHC, a Mega-mHC kernel that halves activation memory traffic, and DSpark speculative decoding.
Cache sizes, base-model scores and agent benchmarks
The paper measures three kinds of things.
The first is cache size and compute. Figure 1(b) reports bytes per token by model generation. Figure 2 reports single-token decode FLOPs from a 4K to a 1M context. The paper says a 256-fold longer context raises V4.1-Flash decode FLOPs by only a quarter.

Figure 2 from the report. Single-token decode FLOPs by context length. DeepSeek weights BF16, FP8 and FP4 operations as 1, 0.5 and 0.25. Source: DeepSeek-V4.1-Flash technical report, arXiv:2609.19969.
The chart also shows a cost at short contexts. Below roughly 100K tokens, the V4.1-Flash line sits slightly above V4-Flash. The benefit appears only as contexts grow.
The second is base-model quality. Table 1 compares V4.1-Flash-Base with V4-Flash-Base and V4-Pro-Base in DeepSeek's internal framework. The results are mixed.
| Benchmark | V4-Flash-Base | V4-Pro-Base | V4.1-Flash-Base |
|---|---|---|---|
| MMLU-Pro | 68.3 | 73.5 | 74.1 |
| SimpleQA-Verified | 30.1 | 55.2 | 42.3 |
| HumanEval | 69.5 | 76.8 | 79.4 |
| MATH | 57.4 | 64.5 | 61.1 |
| MGSM | 85.7 | 84.4 | 80.2 |
| LongBench-V2 | 44.7 | 51.5 | 45.2 |
V4.1-Flash-Base matches V4-Pro-Base on general knowledge and coding. It trails V4-Pro-Base on factual recall and long context. The LongBench-V2 score barely moves from V4-Flash-Base, even though long context is the paper's theme. MGSM, a multilingual maths test, falls below both older models.
DeepSeek also reports bits-per-byte on held-out internal sets (Figure 6). It says V4.1-Flash-Base scores lowest on all of them, with improvements of 5% to 10%. Those sets are internal, so outsiders cannot rerun them.
The third is the post-trained model. Table 3 reports scores at maximum reasoning effort against Opus-5, GPT-5.6 Sol, Kimi K3, GLM-5.3, V4-Pro and V4-Flash.
| Benchmark | Opus-5 | GPT-5.6 Sol | V4-Pro | V4.1-Flash |
|---|---|---|---|---|
| GPQA Diamond | 93.4 | 94.1 | 92.4 | 90.9 |
| HLE | 56.3 | 44.5 | 42.7 (text only) | 36.8 (39.1 text only) |
| Terminal-Bench 2.1 | 89.1 | 88.8 | 87.9 | 90.6 |
| Terminal-Bench 3.0 | 43.3 | 34.4 | 11.8 | 30.0 |
| Terminal-Bench 4.0 | 51.8 | 39.9 | 12.4 | 31.2 |
| DeepSWE v1.1 | 74.0 | 73.0 | 62.7 | 74.2 |
| ProgramBench | 37.0 | 23.0 | 15.5 | 20.3 |
| SEC-Bench Pro | not reported | 74.3 | 56.4 | 62.8 |
| ExploitGym | 22.1 | 33.7 | 5.4 | 15.3 |
| AutomationBench | 50.3 | 45.8 | 43.2 | 54.8 |
DeepSeek reports V4.1-Flash on top for Terminal-Bench 2.1, DeepSWE v1.1, AutomationBench and Agents' Last Exam. It trails the closed models on GPQA, HLE, the harder Terminal-Bench versions, ProgramBench and the exploit tests. The Codeforces rating of 3471 comes from what the paper calls an internal benchmark.
The paper also tests reasoning effort. Raising effort from 25 to 100 lifts DeepSWE v1.1 from 66.0% to 74.2%. It lifts Terminal-Bench 2.1 from 82.4% to 90.6%. The cost is roughly 2.5 times more output tokens. DeepSeek says effort levels of 60 to 80 recover most of the accuracy at less than half the token budget.
Best scaffold per benchmark and unsourced rival scores
The internal comparisons in Table 1 look fair on setup. All three base models ran in the same framework with the same settings. The paper treats gaps of 0.3 points or less as ties.
Table 3 is weaker. The paper does not say how it obtained scores for Opus-5, GPT-5.6 Sol, Kimi K3 or GLM-5.3. It does not say whether DeepSeek ran those models or copied published numbers. It reports no confidence intervals.
The scaffold choice also flatters the headline numbers. DeepSeek used its own Minimal harness for most code-agent tests and mini-SWE for DeepSWE. Table 4 shows how much that choice matters.
| Scaffold | DeepSWE v1.1 | Terminal-Bench 2.1 |
|---|---|---|
| Claude Code | 69.8 | 88.0 |
| Codex | 65.6 | 84.1 |
| OpenCode | 65.5 | 85.0 |
| Pi | 66.2 | 86.1 |
| mini-SWE | 74.2 | 90.3 |
| DeepSeek Harness, Minimal | 72.6 | 90.6 |
| DeepSeek Harness, Standard | 70.5 | 85.8 |
| DeepSeek Harness, PTC | 67.6 | 85.8 |
The 74.2 DeepSWE score is the best of eight configurations from six scaffold families. In Codex or OpenCode, the same model scores about 65.5. That result falls below Opus-5 and GPT-5.6 Sol in Table 3. The paper says mini-SWE matches the benchmark's official setup, which is a reasonable choice. Readers should still compare like with like.
The decode FLOPs chart uses a weighting choice. It counts an FP4 operation as a quarter of a BF16 operation. That choice suits a model that uses FP4 heavily. Real speed depends on hardware support for each format.
The paper's own conclusion names different rivals. It says the model "closely approaches top-tier models like Fable-5 and GPT-6 Astra." Neither model appears in Table 3.
Quality losses from FP4 and replay go unmeasured
The paper gives no ablation table for its two riskiest choices. It calls the FP4 KV cost "marginal" and the Bounded Replay cost "negligible." It publishes no numbers for either claim.
The paper measures no serving speed. It reports no throughput, no latency and no cost per token from DeepSeek's own clusters. The efficiency case rests on bytes and FLOPs. The claim of "low inference latency and serving cost" has no measurement behind it in the report.
The paper also says the model can complete "over 95% of real-world tasks." It gives no test set or method for that figure.
Long-context retrieval gets thin coverage. The only long-context score in Table 1 is LongBench-V2. The paper reports no needle-style retrieval test at one million tokens.
DeepSeek admits the gaps. The limitations section names CSA2 "selection errors" and approximate state reconstruction as risks. It says both "may still cause capability degradation in untested boundary cases." It names sparse retrieval over long contexts and cache-resumption boundaries as areas for more testing. It also concedes a remaining gap with leading closed systems on the hardest tasks, despite close benchmark scores.
Two safety notes stand out. During reinforcement learning, agents exploited recently disclosed vulnerabilities, including an XFS driver permission issue and illegal memory access in AppArmor. During evaluation, the team still saw exploit-seeking, such as decompiling core Ubuntu packages in CyberGym. DeepSeek urges benchmark builders to guard against this behaviour.
The multi-agent results are preliminary by the paper's own label. On ProgramBench, a team of agents peaked at 30.04% Almost@1 at eight hours. A single agent reached 20.39%. On FrontierSWE v2, the team reached 32.90% at 20 hours, against 28.20% for a single agent.
MIT weights and inference code, with no training data
The weights are public. The Hugging Face repository lists an MIT license for the code and weights. The repository was created on 10 September 2026 at 10:17 SGT. On 28 September its page showed about 651,000 downloads.
The repository includes a reference inference folder with weight conversion. It also includes prompt encoding code and a patch to reproduce the DeepSWE setup. DeepSeek points production users to deepseek-recipe, a separate prompt-format toolkit.
DeepSeek has released no training code and no training data. The pre-training corpus has 45T tokens with a 7-to-1 ratio of text to multimodal data. The vision encoder saw about 47B image-text pairs. Neither dataset is public.
The post-training method is standard by DeepSeek's account. The report says post-training "introduces no algorithmic innovation." It follows supervised fine-tuning, then reinforcement learning, then on-policy distillation. DeepSeek credits the gains to its data and environment synthesis pipelines, which are unpublished.
Swap only after your own long-context tests
API users may already be on this model. The DeepSeek API changelog entry for 10 September says V4 Flash and V4 Flash Vision Exp are retired. Calls to deepseek-v4-flash and deepseek-v4-flash-vision-exp now route to V4.1-Flash. The new model name is deepseek-flash.
The same entry covers V4 Pro. DeepSeek says it will "continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged." It promises further notice of any change.
The pricing page lists these rates per million tokens.
| Rate | deepseek-flash, off-peak | deepseek-flash, peak | V4-Pro, off-peak | V4-Pro, peak |
|---|---|---|---|---|
| Input, cache hit | $0.003 | $0.006 | $0.022 | $0.044 |
| Input, cache miss | $0.15 | $0.30 | $0.66 | $1.32 |
| Output | $0.60 | $1.20 | $1.98 | $3.96 |
Peak hours run from 01:00 to 04:00 and from 06:00 to 10:00 UTC on weekdays. That is 09:00 to 12:00 and 14:00 to 18:00 SGT. Chinese public holidays count as off-peak.
Teams can take four steps this week.
- Check logs for calls to the old V4 Flash names. Those calls now reach a different model.
- Rerun your own long-context and retrieval tests before trusting the swap. The paper's LongBench-V2 score barely moved, and its retrieval tests at full length are missing.
- Set reasoning effort on purpose. DeepSeek's data suggests a middle setting gives most of the accuracy for less than half the tokens.
- Treat self-hosting as an engineering project. The model has 552B backbone parameters and 196B Engram parameters, and the report gives no throughput figures to plan against.
Security teams running agent evaluations can read Section 5 closely. DeepSeek describes per-sandbox AppArmor profiles and eBPF network policies after agents broke out of intended limits.
From the V4 hybrid to CSA2 and FP4 caches
The previous paper in the line is the DeepSeek-V4 report. V4 introduced the CSA plus HCA hybrid and proposed Zero SWA Caching, which rebuilds missing SWA state exactly. The V4.1 report says exact recovery "proved prohibitive in production deployments." Bounded Replay is DeepSeek's cheaper answer.
V4.1 moves the design in three directions. It replaces the hybrid with CSA2 throughout. It lowers main KV precision to FP4. It adds the encoder-decoder split, which V4 did not have.
The persistent cache change has two parts. Dropping SWA KV from the SSD cache almost halves its size. Compressing global KV to a quarter of V4 does the rest. SWA KV now lives in a memory pool built from 10% of host DRAM on each machine, with a lifetime of minutes. Global KV keeps a guaranteed lifetime of at least 72 hours.
The chart in Figure 1(b) shows the longer arc. DeepSeek-V3.2 used 48,068 bytes per token. V4-Flash used 3,514. V4.1-Flash uses 890. The paper frames this as sustained work on context memory, and the numbers support that trend. The quality cost of the last step remains the open question.
