Back to timeline

Availability · September 10, 2026

DeepSeek ships V4.1-Flash with an encoder and a decoder

A 552-billion-parameter open-weights model steps away from the decoder-only architecture and, according to the developer, needs a quarter of the HBM and an eighth of the SSD for its cache.

Why it matters

A frontier open-weights model left the architecture every major line had used since 2020, and the developer stated the consequence in memory rather than in scores.

The announcement is dated 10 September 2026. DeepSeek-V4.1-Flash is a mixture of experts with 552 billion backbone parameters and a new causal encoder-decoder architecture: 40 layers, 20 in the encoder and 20 in the decoder. It activates 8 billion parameters per token during prefill and 16 billion during decode. The key-value cache needs a quarter of the HBM and an eighth of the SSD storage of the previous generation. The model is natively multimodal and served through the API under the identifier deepseek-flash; new pricing took effect at 04:00 UTC on 10 September 2026. The weights are published on Hugging Face under the MIT licence. The developer reports that tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed and total runtime, and on that basis V4-Pro is being retired in stages.

Event record

Event date
September 10, 2026
Timeline date
Event date
Verification
Sources gathered automatically · September 19, 2026
Lines
ID
evt-0412

The day of the announcement in DeepSeek's API news; the new pricing took effect at 04:00 UTC the same day.

Sources

Related events