金色财经|Sep 29, 2026 08:15
[NVIDIA Labs: Optimizing Kimi Linear Attention with Kernel-Designed Agents, Achieving a 2.96x Speedup]
According to a report by Golden Finance on September 27, NVIDIA Labs' blog 'KDA²' revealed that its kernel-designed agent (KDAgent) is now capable of autonomously optimizing Moonshot AI's Kimi Delta Attention (KDAttn) operator. On NVIDIA B300 with 8,192 tokens, compared to the official FlashKDA baseline: the TIRx version achieved a geometric mean speedup of 2.96x, the CAKE-PTX version 2.94x, and the CuTe-DSL version 2.85x. Additionally, the new kernel reduced the relative error of the final state for 8k tokens from FlashKDA's 3.45% to 0.22% (CuTe) and 0.29% (TIRx), approximately one-tenth of the original. The kernel has been open-sourced.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink