DeepSeek launching v4.1 flash cheaper and more capable than v4 pro

ThinkingNews Desk · how this was written
DeepSeek will launch its new V4.1 Flash model on September 10 2026 Beijing Time, positioning it as cheaper and more capable than the current V4 Pro across performance, cost, speed, and task completion metrics. The compact model retains a 552-billion-parameter backbone while supporting a 1-million-token context window and is built on a new causal encoder-decoder architecture. After launch, all V4 Pro requests will be redirected to V4.1 Flash and billed at Flash rates.
Written from all 3 reports below, not from any single one.
How it was reported
- Hacker News·DeepSeek launching v4.1 flash cheaper and more capable than v4 pro
DeepSeek will launch the V4.1 Flash model on September 10 2026 (Beijing Time), positioning it as cheaper and more capable than the current V4 Pro across performance, cost, speed, and task completion metrics. After launch, all V4 Pro requests will be redirected to V4.1 Flash and billed at Flash rates, with off-peak pricing set at $0.003 per input cache hit, $0.15 per miss, and $0.6 per output, while peak-hour rates will be double.
- Hacker News·DeepSeek v4.1 Flash
- TechMeme·DeepSeek debuts DeepSeek-V4.1-Flash, its smallest model built on a new Causal Encoder-Decoder architecture, with 552B backbone parameters and 1M-token context (Reuters)
DeepSeek introduced DeepSeek-V4.1-Flash, a compact model built on a new Causal Encoder-Decoder architecture that retains a 552-billion-parameter backbone while supporting a 1 million-token context window. The design targets efficiency without sacrificing the large-scale reasoning capabilities of its larger predecessors.
- The Next Web·DeepSeek launches V4.1-Flash and retires V4-Pro, its flagship model
DeepSeek’s V4.1-Flash is a 552-billion-parameter causal encoder-decoder that activates 8 billion parameters per input token and 16 billion per output token, supports up to a 1 million-token context, and includes native image understanding while using 890 bytes per token for KV cache. The company will retire V4-Pro on 14 September, redirecting all requests to V4.1-Flash at lower rates. Benchmark scores are 74.2 on DeepSWE v1.1, 88.1 on CyberGym, and 36.8 on Humanity’s Last Exam.
Related stories
- Three reasons why DeepSeek’s new model V4 matters8 outlets
- DeepSeek unveils an experimental version of its V4 Flash model that can understand visual prompts, saying it nears the performance of Anthropic's Opus 4.8 (Bloomberg)3 outlets
- DeepSeek releases DeepSeek Harness under the MIT license in developer preview, touting a design where "every capability is a plugin" that can be swapped out (Carl Franzen/VentureBeat)3 outlets
- DeepSeek made AI cheap. Now it is raising $8bn and buying robots3 outlets