rick awsb ($people, $people)
rick awsb ($people, $people)|Sep 07, 2026 13:30
Astra's pre-training used over 100,000 GB200 cards, and according to Jensen Huang, there are another 400,000 cards about to come online. But compared to the total compute power OpenAI possesses, this is still just a small fraction. 100,000 GB200 cards only accounted for about 10% of OpenAI's total compute power at the time. However, Total Compute ≠ Frontier Training Compute. A model company might have compute power equivalent to millions of GPUs, but these cards are distributed across different data centers, cloud providers, chip generations, and are tasked with supporting ChatGPT, API, Agent, RL, evaluations, synthetic data, and more. What truly determines the upper limit for training a large model is: Maximum Tightly-Coupled Training Compute —— How many cards can work together like "one machine" in a single training session. Training large models requires frequent synchronization of parameters, gradients, activations, and MoE token routing. No matter how fast GPUs compute, if a lot of time is spent waiting on the network, the actual training efficiency will still be very low. This also explains why Astra might only account for 5%-10% of OpenAI's total compute power at the time, yet could have consumed a significant proportion of the maximum unified training resources. So, when Jensen Huang talks about "400,000 GPUs coming online next," the real key is how many of them can actually be integrated into the same training fabric. OpenAI leads in this area because, in addition to their self-built infrastructure, they also have access to Microsoft's rented clusters. In a competitive environment where compute equals models, Anthropic's conservatism at the time now seems likely to be a fatal mistake.
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads