The Frontier · Tags
#KV cache
1 published article
DeepSeek's 17 September V4.1-Flash report cuts global KV cache to 890 bytes per token and leaves the quality cost unmeasured
The 51-page arXiv paper pairs FP4 KV caching, cross-layer cache reuse and an encoder-decoder split. It reports no serving speed and no ablation for its two riskiest shortcuts.