律动BlockBeats
律动BlockBeats|Aug 05, 2026 04:39
[Cursor Reorganizes GPU Pipeline, Speeds Up MoE Model Training by 41%] According to monitoring by Beating Insights, Cursor has open-sourced Mixture-of-Kittens (MoK) to accelerate the training of large-scale MoE models. It integrates previously separate GPU data transfer and computation into a single kernel (a low-level program running directly on the GPU). MoE splits the model into numerous 'experts,' only a portion of which are called upon during each operation. These experts are distributed across different GPUs, requiring frequent data transfers, which can sometimes consume more than half of the training time. MoK enables GPUs to transfer data and perform computations simultaneously, reducing idle time. In real-world training on 512 GB300 GPUs, overall throughput increased by 41%. In isolated tests of the MoE layer, it was up to 2.37 times faster than publicly available solutions. MoK has already been deployed in Cursor's Composer training across tens of thousands of GPUs and is open-sourced under the Apache 2.0 license. Currently, it only supports NVIDIA Blackwell GPUs and is primarily targeted at organizations with GB200 and GB300 NVL72 clusters. [Original Link]
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads