律动BlockBeats|Sep 10, 2026 07:25
DeepSeek breaks the impossible triangle of big models: stronger, faster, and cheaper
Dynamic Beating AI News: DeepSeek V4.1 Flash has almost completely redone the entire architecture, aiming to enhance its capabilities, increase speed, and reduce costs at the same time. Stronger: The model has 552 billion backbone parameters and is also equipped with an embedded 1960 billion parameter Engram conditional memory. Pre training used 45T multimodal tokens, while post training extensively incorporated real agent tasks, tool environments, and failure cases. DeepSWE v1.1 has reached 74.2%, surpassing Claude Opus 5 and GPT-5.6 Sol. Faster: The new CED architecture splits the 40 layer model into front and back halves. When reading Prompt, each Token only activates 8 billion parameters, and when generated, it only activates 16 billion. In addition, with CSA2 cross layer multiplexing and DSpark speculative decoding, the context is pulled from 4K to 1M, the length is increased by 256 times, and the decoding computation of a single token only increases by about a quarter. Cheaper: DeepSeek continues to push KV Cache into the dead. Change the main cache to FP4 and add cross layer reuse, leaving only 890 bytes of global KV for each Token, which is about 1/4 of V4 Flash; The cache stored in SSD or memory for a long time is further reduced to about 1/8. V4.1 Flash does not simply equate 'stronger' with 'counting more'. The model scale continues to grow, but only a small part of the parameters are adjusted each time; The context continues to elongate, but the cache is compressed smaller. The ability has improved, but the speed and cost have not been dragged down. [Original link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink