PANews|Sep 18, 2026 05:23
[Zhipu Launches GLM-5.3-FlashX, Achieving Maximum Inference Speed of 200 Tokens/s]
According to Zhipu's official WeChat account, the company has launched GLM-5.3-FlashX, with its API and experience center now open. The official statement claims a maximum inference speed of 200 tokens/s, achieved through inference computing power provided by 100,000 domestically produced chips, along with additional investments in Infra and inference optimization. The base model, GLM-5.3-Flash, was open-sourced on August 26, featuring 320 billion total parameters, 18 billion active parameters, and a context length of 1 million tokens. It was previously anonymously tested under the name 'Ox Alpha.'
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink