qinbafrank|Aug 07, 2026 07:02
Regarding the reduction of HBM on Nvidia Rubin Ultra, previously here is https://(x.com)/qinbufark/status/2083890594796257652? s=46&t=k6rimWsEbo2D2tXolYcM-A、 Today TheInformation also reported, let's talk about:
1. The mainstream SKUs previewed by Nvidia to major customers have shifted towards lower memory (such as 192GB 8-Hi HBM4, even lower than the standard Rubin's 288GB 12 Hi), while maintaining similar peak computing power (about 35 PFLOPS), with mainstream power consumption controlled at around 1800W.
Previously, here was https://(x.com)/qinbufark/status/208391794731578909? S=46&t=k6rimWSEbo2D2TXolYcM-A As we discussed, if the HBM configuration of the Ultra version is lower than that of the Vera Rubin standard version, the advantages of the standard version (Vera Rubin) will become apparent: it has entered mass production and will be delivered on a large scale to 8 CSPs including AWS, Google, Microsoft, Oracle, CoreWeave, etc. in July. The memory specifications are more 'normal' (288GB HBM4 level), and the unit price and system complexity are more controllable, making it suitable for rapid expansion. If Ultra further reduces its memory allocation (even lower than the standard version) and trains/infers large memory sensitive models, the effective capacity of a single card will decrease, and CSP will naturally tend to use more standard versions to stack first.
2. The core reason for the reduction in allocation is still insufficient supply rather than demand. TrendForce points out that overall DRAM supply will remain tight in 2027, with limited space allocated to HBM wafers, and uncertainty in the validation, yield, and production progress of HBM4E (especially higher layers such as 12 Hi/16 Hi).
The real significant reduction in production capacity will not be achieved until the second half of 2027 to 2028 (new factories/expansions such as SK Hynix and Samsung are gradually released, and SK Hynix plans to double its wafer production capacity in 5 years, but the implementation of new production capacity lags behind)
3. Reducing allocation directly drives up the demand for "more GPUs": lower HBM means that the model may not be able to fit or require more splitting/pipeline, ultimately requiring more cards to achieve the same effective computing power or throughput. This will increase the total GPU procurement volume and also raise the system TCO.
So it is estimated that the standard version will still be the most purchased by CSP cloud factories in the future.
4. Reducing allocation also benefits interconnection
If CSP buys more standard or low-end Ultra, the interconnection bandwidth, latency, and software scheduling (NVLink domain expansion, collective communication) between racks/cabinets will be more tight. Nvidia is also strengthening this aspect (higher generation NVSwitch, CPO optics, etc.).
The shortage of HBM has shifted the pressure from the "memory" part to "interconnection+system software+power/cooling". CSP's procurement strategy will place more emphasis on "how many cards can be efficiently interconnected" rather than how strong the single card specifications are.
summary
The shortage of HBM is structural (at least until 2027-2028), and it is a rational choice for Dazhi to reduce its allocation and replace its shipment volume. This will make standard Vera Rubin the "main force" of CSP from 2026-2027, while Ultra will play a more "large-scale interconnection optimization version". The importance of interconnectivity (NVLink ecosystem) will be further amplified and become the key to differentiation in the next stage.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink