
Written by: Chuangyebang
Recently, a reader sent a notice from Zhipu.
The old plan that had been used for several months is about to end.
The previous "unlimited weekly quota" GLM Coding Plan will stop auto-renewal, and Zhipu will compensate old users with two months of the new plan.
Several months ago, the AI programming market was focused on who had the cheaper plan; now, everyone is competing on another aspect:
How many Tokens can you actually use.
On September 2, Zhipu officially opened its flagship store on Tmall, listing the GLM Coding Plan. The Lite version costs 118 yuan per month, the Pro version costs 538 yuan per month, and the Max version costs 1078 yuan per month, corresponding to different point allowances; the products are based on GLM-5.3 and are compatible with over 20 mainstream Agents like ZCode, Claude Code, and Codex.

This means that something that previously sounded very "technical" is beginning to appear on e-commerce shelves:
The usage allowance for AI can now be directly purchased like mobile data.
But what is really worth noting here is not the fact that “Zhipu sells Tokens on Tmall” itself.
Rather, it's a somewhat counterintuitive change:
Tokens are becoming cheaper, but AI companies are starting to care more about how many Tokens you use.
Everyone is starting to limit
In the past few months, there has been a very interesting shift in the AI industry.
Kimi had previously suspended subscriptions for C-end new users due to computing power pressure. On July 19, the announcement from the Dark Side of the Moon stated that after the Kimi K3 went live, the request volume greatly exceeded expectations, approaching the current cluster capacity limit, thus suspending C-end new user subscriptions to prioritize computational power for already paying users.

Alibaba Cloud is also adjusting its Coding Plan. Its Lite package stopped new purchases on March 20, and stopped renewals and upgrades on April 13.
Tencent Cloud adjusted the prices of some mixed models in March. Among them, the input price of HY2.0 Instruct increased from 0.0008 yuan per thousand Tokens to 0.004505 yuan, a rise of over 460%.
Why has AI suddenly stopped being "generous"?
In the past 20 years, the biggest lesson from Internet products has been that after digitization, marginal costs will fall increasingly low.
Mobile data is the most typical example. Once the base stations, fiber optics, and network infrastructure are built, the marginal cost of transmitting an additional GB of data is very low. As infrastructure scales up, data prices plummet.
So many would naturally apply this logic to AI—chips are getting stronger, models are becoming more efficient, and the reasoning cost per Token is getting lower, thus AI should be getting cheaper.
In fact, the first half of this conclusion is completely correct.
Zhipu itself disclosed that as the model architecture and reasoning infrastructure continue to optimize, its unit Token reasoning cost has dropped by 80% since the beginning of the year. The company also stated that it has achieved low-cost reasoning at the scale of 100,000 domestic chips.
The problem is that while Tokens have become cheaper, the rate of Token consumption is growing faster. This is the real contradiction in today’s AI industry.
Tokens are not ordinary data traffic.
Behind a Token are the computing resources consumed by model inference, which include GPUs or AI acceleration chips, high-bandwidth memory, networking, storage, and the electricity and cooling in data centers.
Tokens are more like a unit of AI computing power being measured and traded.
When a model is just chatting with you, a single invocation may only require several hundred or thousands of Tokens. But when you ask AI to complete a complex task for you, the consumption is completely different.
From “selling models” to “selling calls”
Zhipu's latest semi-annual report has recorded this change in its financial data.
In the first half of 2026, Zhipu achieved revenue of 954 million yuan, a year-on-year increase of 399.7%. In the first half of the year, Zhipu's MaaS open platform and API business revenue reached 825 million yuan, a year-on-year increase of 2735.7%, accounting for 86.5% of total revenue.

In the same period last year, this proportion was only 15.2%. Meanwhile, the proportion of localized deployment revenue has dropped from 84.8% to 13.5%.
In one year, Zhipu’s business model has almost completed a shift.
In the past, customers bought models. The models were deployed on the companies' own servers for customization, integration, and delivery, often linking a single project to a single revenue.
Now, more and more customers are directly invoking cloud models. Each invocation incurs a fee; the more they use, the more they pay. This is a typical MaaS business.
Moreover, this change in Zhipu is not simply "lower prices for greater scale." On the contrary, it has resulted in an interesting combination:
The volume of Token calls has grown by over 40 times since the beginning of the year; the average price of APIs has risen approximately 101%; the unit Token reasoning cost has decreased by 80%; and the gross margin of the MaaS business has increased to 24.6%.
Costs are declining, revenues are increasing, users are using more, and the average price is actually increasing. AI commercialization may be breaking away from the initial simple “price war.”
In the early stages, the model's capabilities were insufficient, and platforms could only attract users with low prices or even for free.
But when the models become capable of truly completing code generation, tool invocation, continuous execution, and other complex tasks, users are no longer just buying “a model,” but the ability of the model to do work for them.
This is also why Zhipu's management summarizes the business model into a very interesting path: sell models → sell calls → sell subscriptions → sell end-to-end task results.
These four steps correspond to the four stages of AI commercialization.
What Zhipu really wants to sell is not Tokens
In the past, it was the Chatbot model: you asked a question, it answered. A single interaction may only require several hundred or thousands of Tokens.
Now it’s the Agent model: you say “Help me create a website,” it may need to understand the requirements, search for information, write code, debug tools, run tests, find bugs, modify, and test again… A task may consume dozens of thousands, hundreds of thousands, or even more Tokens.
Moving forward is the Multi-Agent model: one Agent plans, another searches, another writes code, another tests, and another reviews. The demand for Tokens continues to expand.
In Q1 2026, the global weekly Token usage exploded by 250% to 22.7 trillion. The Token invocation volume on the OpenAI platform surged from about 6 billion per minute in October 2025 to 15 billion per minute by the end of March 2026, an increase of 150% in less than half a year. Morgan Stanley reports show that top large language models are experiencing a “non-linear capability leap,” as AI's explosive growth encounters systematic supply bottlenecks.
Thus, what is really happening is not that Tokens have become more expensive, but that everyone is starting to need more and more Tokens. This may be more alarming than simply price hikes.
There is a classic phenomenon in economics called the “Jevons Paradox”:
When the efficiency of using a resource improves and the unit cost decreases, people may not reduce their usage; on the contrary, they may increase usage significantly because it has become cheaper.
AI is experiencing a similar situation.
The cheaper Tokens are, the more people dare to let AI do more things; the smarter AI becomes, the more willing people are to entrust it with more complex tasks.
Thus, Token consumption has formed a new growth flywheel.
Zhipu selling Tokens on Tmall appears to be selling AI quotas. But from the semi-annual report, what it is truly doing is turning model capabilities into continuously billable cloud services.
In 2025, Zhipu’s main income was still localized deployment; by the first half of 2026, MaaS/API accounted for 86.5% of income. The company summarizes this evolution path as: “sell models → sell calls → sell subscriptions → sell end-to-end task results.”
Zhipu’s board secretary, Xiao Lei, candidly stated in the earnings call: “The shift in revenue composition is more noteworthy compared to income growth rate.”
This shift indicates a fundamental change in the business model—from one-time project delivery to continuous service income. And the pricing logic will change accordingly: from “how much is it per million Tokens” to “how much for completing a task.” Zhipu’s chairman, Liu Debing, mentioned on the call: “Leading a specific ranking will soon be equaled, what’s truly important is who can deliver higher intelligence at lower costs continuously.”
The trend for the future may be as follows:
The basic Tokens will become cheaper and cheaper, but advanced intelligence will become more and more expensive.
Just like cloud computing: servers are getting cheaper, but the money businesses spend on cloud services hasn’t disappeared because they are using more.
AI is similar. A decline in single Token prices does not equal a decline in AI usage costs. Because the stronger the models become, the more complex tasks people assign to them. From asking a question to completing a task, to managing an entire workflow, Token consumption will continue to expand.
A phrase in Zhipu’s semi-annual report is worth repeating: “Every time the model makes a leap in capability, what the customer buys gets closer to the result, and our revenue structure is rewritten once.”
Therefore, what may be truly worth paying attention to in the future is not “how many Tokens can be bought for 1 yuan,” but how much AI intelligence a person or a company actually consumes in a month.
Putting Tokens on Tmall may only be the first step towards broadening the “AI metering economy.”
免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。