深潮TechFlow|Sep 04, 2026 12:59
[GLM-5.3-Flash Tops B.AI Model Invocation Rankings, Cumulative Throughput Exceeds 24.1 Trillion Tokens]
DeepFlow Tech News, September 4: The GLM-5.3-Flash Niulai model has become the most invoked and popular model on the B.AI platform, with cumulative token throughput surpassing 24.1 trillion. As the first native multimodal model in the GLM-5 series, GLM-5.3-Flash features 320B total parameters and 18B active parameters, utilizing a hybrid architecture that combines sparse and linear attention. It supports a 1M ultra-long context, balancing rapid response, powerful reasoning, and high cost-effectiveness. Starting today, developers can continue to invoke this model for free via the B.AI platform, covering diverse scenarios such as high-frequency APIs, code generation, complex agents, and ultra-long document processing. Experience it now: chat.b.ai/chat
Share To
HotFlash
APP
X
Telegram
CopyLink