PANews
PANews|Sep 18, 2026 05:23
[Zhipu Launches GLM-5.3-FlashX, Achieving Maximum Inference Speed of 200 Tokens/s] According to Zhipu's official WeChat account, the company has launched GLM-5.3-FlashX, with its API and experience center now open. The official statement claims a maximum inference speed of 200 tokens/s, achieved through inference computing power provided by 100,000 domestically produced chips, along with additional investments in Infra and inference optimization. The base model, GLM-5.3-Flash, was open-sourced on August 26, featuring 320 billion total parameters, 18 billion active parameters, and a context length of 1 million tokens. It was previously anonymously tested under the name 'Ox Alpha.'
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads