律动BlockBeats
律动BlockBeats|Sep 29, 2026 08:57
[Databricks Sweeps NVIDIA Kernel Benchmark with AI: First Place in Four Tracks Across 235 Tests, Token Cost Around $70,000] Dongcha Beating AI Newsflash: The Databricks team claims they placed GPT-6 Astra and Opus 5 into an automated optimization loop, enabling the models to continuously rewrite, test, and improve GPU kernels. Ultimately, they secured first place in all four tracks of NVIDIA's SOL-ExecBench. This benchmark includes 235 GPU kernel tasks, covering basic operators, complex fused operators, low-precision computation, and real-world large model inference. The team stated that the total token cost for the process was approximately $70,000. This system is built on KDA and Humanize, allowing AI to write kernels, run correctness and performance tests, and then iteratively modify based on the results. The KDA team recently applied a similar method to Kimi Delta Attention, achieving 2.96 times the speed of the official FlashKDA on NVIDIA B300. Both achievements are different applications of the same agent-based kernel optimization framework. Kimi demonstrates that a high-complexity attention kernel can be rewritten by AI, while Databricks extended this approach to hundreds of different kernels this time. [Original Link]
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads