金色财经
金色财经|Sep 29, 2026 08:15
[NVIDIA Labs: Optimizing Kimi Linear Attention with Kernel-Designed Agents, Achieving a 2.96x Speedup] According to a report by Golden Finance on September 27, NVIDIA Labs' blog 'KDA²' revealed that its kernel-designed agent (KDAgent) is now capable of autonomously optimizing Moonshot AI's Kimi Delta Attention (KDAttn) operator. On NVIDIA B300 with 8,192 tokens, compared to the official FlashKDA baseline: the TIRx version achieved a geometric mean speedup of 2.96x, the CAKE-PTX version 2.94x, and the CuTe-DSL version 2.85x. Additionally, the new kernel reduced the relative error of the final state for 8k tokens from FlashKDA's 3.45% to 0.22% (CuTe) and 0.29% (TIRx), approximately one-tenth of the original. The kernel has been open-sourced.
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads