律动BlockBeats|Jul 19, 2026 02:21
[SemiAnalysis: Kimi K3 Significantly Reduces KV Transmission Bandwidth, but AI Network Demand Will Not Shrink as a Result]
BlockBeats News, July 19 — Independent semiconductor and AI research institution SemiAnalysis published an article stating that although approximately three-quarters of Kimi K3's network layers adopt KDA, which can reduce KV cache transmission bandwidth by up to 10 times compared to a full global attention model, this does not imply that the market size for AI network switches will shrink significantly.
Kimi K3 boasts 2.8 trillion parameters, and even with MXFP4, each forward computation still requires approximately 1.5TB of HBM bandwidth. To achieve profitable deployment while maintaining reasonable interaction speeds, it is still necessary to connect a large number of chips through high-bandwidth networks like GB300 NVL72 and rely on WideEP extension services. WideEP distributes 896 expert models across multiple GPUs and performs Token distribution and result merging twice per layer and per forward computation, requiring over 120 executions for a single forward computation.
In contrast, KV cache transmission between pre-filling and decoding occurs only once per conversation round. Therefore, the bandwidth savings achieved by KDA may be far smaller than the expanded network demand brought about by large-scale expert models. SemiAnalysis believes that more efficient attention mechanisms could further push context lengths from 1 million Tokens to over 5 million Tokens. According to Jevons Paradox, efficiency improvements may expand AI usage scale, thereby further increasing network demand. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink