律动BlockBeats|8月 31, 2026 11:34
[Zhipu Token Inference Costs Down 80% Since Early This Year: 100,000 Domestic Chips Running Large Models]
Beating AI News Flash: Zhipu has disclosed that the company has now achieved large-scale, low-cost inference using 100,000 domestic chips, with per-token inference costs reduced by 80% compared to the beginning of the year. This batch of domestic computing power has already undergone real traffic stress testing. Before the release of GLM-5.3-Flash, it was anonymously launched on OpenRouter and OpenCode under the codename Ox-Alpha, with token calls reaching 62 trillion. After the official release, all online traffic continues to be supported by 100,000 domestic chips. [Original Link]
Share To
HotFlash
APP
X
Telegram
CopyLink