DeepSeek V4.1 Flash Officially Released: 552B New Architecture, Only 8B Activated for Input
律动BlockBeats|Sep 10, 2026 06:11
Beating AI News Flash: DeepSeek has officially released V4.1 Flash. The new model features a total of 552B (552 billion) parameters and adopts a brand-new Causal-Encoder-Decoder architecture with native support for image understanding. It is also the smallest model in DeepSeek's new architecture series. This architecture separates the processing of input and output. When reading input, only 8B (8 billion) parameters are activated, and 16B (16 billion) are activated when generating responses. DeepSeek claims this design reduces inference costs even for large models with 552B parameters.
After new pretraining and larger-scale reinforcement learning, V4.1 Flash has surpassed V4 Pro in official benchmarks. Cache usage has also been significantly compressed. Compared to the previous generation, V4.1 Flash reduces HBM memory requirements to 1/4 and SSD storage requirements to 1/8, primarily to lower costs for long-context and repeated Agent calls.
V4.1 Flash is now available via API, and the model can be accessed by using the name `deepseek-flash`. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink