Hacker News
DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression
DeepSeek-V4.1 Flash compresses KVCache by 4×, using head-count, block-based, and cross-layer compression techniques, and optimizes prefill computation so only 20 of the 40 layers activate, reducing activated parameters to 8 B for prefill and 16 B for decode. The model also adopts FP4 precision for KVCache and refines sparse-attention indexing to lower interconnect bandwidth and storage pressure on HBM and SSD. These changes enable ultra-long-context processing for long-horizon agent workflows while maintaining task-completion quality.