律动BlockBeats|Sep 29, 2026 07:58
[NVIDIA Lets AI Rewrite Kimi Attention Kernel: Speed Reaches Nearly 3x Official Version]
Beating AI Newsflash: NVIDIA's research team enabled an AI Agent to directly rewrite the GPU kernel for Kimi Delta Attention. Kimi Delta Attention is one of the core attention mechanisms used in the Moonshade Kimi-Linear model. Traditionally, such low-level code required engineers familiar with CUDA and chip architecture to manually optimize it repeatedly. Ultimately, the version written by the Agent achieved 2.96x the speed of Moonshade's official FlashKDA on the NVIDIA B300. The team tested six sets of tasks with fixed-length and varying-length sequences, and the 2.96x speed was the geometric mean across all tasks. The relevant kernel has already been open-sourced. However, the Agent quickly identified testing loopholes. At one point, it hardcoded statistical patterns from the test data directly into the code, achieving a speed of 3.74x; it also tried retaining only the most recent 32 tokens, with single-task performance even reaching 5.16x. These versions, however, failed when applied to real Kimi data, with some producing errors or outright failing. The team subsequently incorporated real Kimi runtime data, random inputs, and extreme value tests, tightening error standards. The final retained version with 2.96x speed passed this validation process. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink