After the semiconductor stocks plummeted by 40%, re-examine the fundamentals of computing power demand.

CN
5 hours ago
When the backlog of computing orders from ultra-large-scale cloud vendors surged from $500 billion to $2 trillion, stronger evidence is needed to claim it is a bubble.

Author: Kerman Kohli

Translated by: Deep Tide TechFlow

Deep Tide Introduction: After semiconductor stocks plummeted 30-40%, many declared the AI bubble burst. However, this judgment is based on a fatal assumption: that the demand for computing power is limited. From government military, scientific research to corporate products and personal applications, all groups are vying for computing power, and the price ceiling each group is willing to pay is different. More importantly, AI has a characteristic that other infrastructures do not have — recursive demand: computing power itself will generate more demand for computing power. When the backlog of computing orders from ultra-large-scale cloud vendors surged from $500 billion to $2 trillion, stronger evidence is needed to claim it is a bubble.

Core Question: Infinite Demand or Limited Demand

According to the timing of this article's release, semiconductor and other AI/momentum-related stocks have already fallen 30-40% from historical highs.

Many are eager to call this the peak of semiconductors/AI/memory and celebrate their victory for not participating.

They are probably celebrating too early.

In my view, the semiconductor/AI investment argument boils down to this one question:

"Do you think the demand for computing power is limited or infinite?"

In conversations, I've seen too many people trapped in their localized experiences of corporate usage/adoption and generalizing those to a wider market. I believe the argument that corporate adoption takes longer may indeed hold true.

However, this also creates a reverse incentive, allowing smaller, more AI-native companies to outperform them with fewer employees, because AI, if used properly, can be far cheaper than expanding with human labor.

In any case, let’s broaden our perspective and stop viewing AI capital expenditures merely as corporate demand. At a high level, AI deployment is about bringing computing power online.

Computing power, in turn, can and will be utilized by the following groups:

  • Governments for military and defense purposes
  • Scientists for medicine and other cutting-edge research
  • Businesses to build new products and expand without labor constraints
  • Individuals empowered to do more (programming, designing, creating, questioning)

These use cases and buyers are each willing to pay different prices for computing power. While some may drop out due to computing power costs, it is unwise to believe that no one else will have a higher budget. For some, spending on computing power is non-negotiable because it’s a brutal competition. Examples include sovereign governments and ultra-large-scale cloud vendors. For others, computing power is a substitute for labor costs and still far cheaper (without legal overhead, management time, etc.).

Believing we are "overbuilding" or "building beyond capacity" implies that the aforementioned groups have reached a final state and are satisfied with the status quo. Simply put, it believes:

  • Governments think they do not need smarter weapons and defense capabilities
  • Scientists are happy with the amount of research completed
  • Businesses think they’ve done enough work on product lines and do not want to grow further
  • Individuals have reached the final state of curiosity and do not want to do more

If you truly believe any of the above, then you are correct to say that AI is a bubble and that there will be overinvestment.

Each group is willing/able to pay different prices for computing power, but the market will organize around demand to ensure quality/price can be met. Believing intelligence is too expensive for all groups is a misunderstanding, as those able to profitably harness intelligence will continue to drive demand for it.

However, if you believe that humanity and the aforementioned groups will never be satisfied, then you must believe that the demand for computing power is infinite. We are currently in a massive race around computing power, and few people realize this.

Follow-up Question: What is the Return on Investment

After spending more time in the market, this is the investors' greatest ongoing concern regarding computing power construction. Is the trend of ultra-large-scale cloud vendors spending all their free cash flow excessive, are they gambling with their future?

Beyond the capital expenditures of ultra-large-scale cloud vendors, there is also significant concern about what revenue and profits large labs are achieving? The emergence of new open-source models threatens to extract value from the cutting-edge models, complicating the situation further.

I will take time to discuss each point, but let's begin with the capital expenditures of ultra-large-scale cloud vendors.

For those who believe ultra-large-scale cloud vendors have miscalculated, I need you to understand this is far from true: public cloud services are outrageously expensive, and they know how to squeeze every penny from you. They have convinced an entire generation of companies that you cannot scale without them.

In return, they have marked the cost of normal non-CPU computing power up 10-20 times. Additionally, you also have to pay for logging, data transfer (outbound), and another five services just to complete basic work.

This game works because they have tied you into their system. Bandwidth within the GCP/AWS kingdom is very cheap, but once you move out, it skyrockets. For many ultra-large-scale cloud vendors, they need customers to remain within their ecosystem, otherwise, there’s a risk of losing business. Lack of sufficient computing power is fatal to their survival. When your customer data and computing power are all with you, saying your GPU runs out is utterly unacceptable and will force them to slowly turn to your competitors. Ultra-large-scale cloud vendors have created an interesting dynamic where they can force customers to pay whatever price they want, and customers are powerless to resist. Their tenants are so disconnected from bare metal that there is a huge lock-in effect, making transfer a multi-year effort (if possible).

A simple example to illustrate how crazy they are in this respect. A few months ago I wrote about how I built this $15,000 machine:

Building an AI Inference Machine

Image: AI inference machine built by the author Source: Kerman Kohli / Substack

You can find the same GPU, with lower specs, on GCP:

https://cloud.google.com/products/compute/pricing/accelerator-optimized

On-demand costs about $3,248 per month, and $1,444 per month for a three-year reservation.

Image: Example of Google Cloud GPU instance pricing. Source: Google Cloud

My machine has only 128GB of DDR5, while Google Cloud’s has 180GB of some memory (they don’t tell you if it’s DDR4 or DDR5, haha).

Quick calculations:

  • On-demand pricing means my machine will pay itself back in 4.6 months
  • The break-even period for a three-year commitment is 10 months

The math for other machines is similar. The H200 cluster (GPU set to release at the end of 2024) will pay itself back in less than 2 years. You can see very similar math everywhere. Of course, this does not account for: land costs, ongoing electricity costs, financing costs, and onsite staff, but it should serve as an illustrative example of how ultra-large-scale cloud vendors know how to price at a massive premium. Of course, spot prices and commitment prices differ again.

What makes this math crazier is that GPUs from five years ago were: a) holding value b) leasing costs were rising!

Hardware is likely to appreciate rather than depreciate. While new chips with higher computational efficiency have been launched, their lack of efficiency will be compensated by rising memory costs.

I conceptualize hardware as a dual-component game where one component (computation) will technically become less valuable but will be offset by another component (memory), which will become more valuable over time.

If this is the case, then their capital expenditures’ return on investment is far higher than anyone remotely expected. Gavin Baker’s tweet summarizes the credit risk about ultra-large-scale cloud vendors well:

Image: Gavin Baker's tweet. Source: X / @GavinSBaker

Now you might say, how do we know this demand is sufficient? What I mean is, I can’t model every scenario for every client, but at some point, you need to humbly say that people queuing up to pay is the strongest signal that you believe they are rational actors spending on positive ROI efforts.

Looking at it from this angle, the backlog of orders has grown from $500 billion at the beginning of 2025 to far beyond $2 trillion in just 1.5 years. When this is driven by customers, claiming that all this is fake/not positive ROI becomes strained. Now, the counterargument is that labs account for a significant portion of this, but that view is incorrect. According to EpochAI, cutting-edge labs account for part of it but are not the entirety of global computing power demand.

Image: Computing power orders backlog has grown from $500 billion to $2 trillion. Source: EpochAI

Image: Composition of global computing power demand. Source: EpochAI

Regardless of what you believe, the fact of a $2 trillion backlog order should indicate something. The notion that trillions in spending does not reflect a structural shift but rather a bubble is an interesting perspective.

Many investors enjoy reasoning through analogy, comparing it to past infrastructure projects like internet construction or railroads. I understand the logic here, but it overlooks a key characteristic of AI: recursive demand.

For railroads or the internet, you need more people to adopt the technology and then ensure that everyone’s usage of that technology reaches a limit to guarantee sufficient diffusion within the economy. AI does not have these dynamics. In this race, computing power can generate its own demand for computing power, and an individual’s or organization's limit in computing power is, in effect, infinite. If you find a useful, positive ROI use case, you can continue to invest in computing power, and it will become a money-making machine.

The tricky part is that different people have very different experiences with AI. Most people in the world use it as a single prompt-answering machine. For people like me, as agent engineering becomes more capable, it is becoming indispensable, allowing me to do more things.

My spending on computing power continues to rise and will keep rising as I find more positively ROI use cases. Regardless of how much revenue AI generates, the cost savings it creates are undeniable, driving the case for most end customers.

Uncertainties: OpenAI / Anthropic

Continuing on the previous point, we can see that the demand backlog is crazy. But how real is the demand from labs (a significant portion of computing power demand)? This is where I think the answer is less clear but still inferable. I want to break this answer into reasoning and training.

If we consider that a new SOTA (state-of-the-art) model costs up to hundreds of millions of dollars, we can say it is an investment asset that generates some useful lifecycle value through reasoning over time (despite having a steep depreciation curve).

As a counteraction, you will see open-source models entering the market and competing for computing power with cutting-edge labs at a cheaper cost. These open-source models may be distilled or not; it does not matter for understanding the dynamics.

So we must question the dynamic of what happens when SOTA models come out? The reality is that not everyone will consistently use them to solve every problem. However, given that they have SOTA capabilities, they can solve problems that current model classes cannot, and you would be willing to pay a premium for that.

You could say that models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are extremely large, and the hardware required to run them far exceeds what any consumer model can do. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding that.

What makes the model more interesting is that Kimi depleted capacity as soon as it opened the floodgates to K3. Of course, there is a model out there, but it still needs someone to serve it. It still has to run on capable hardware. Labs will indeed be forced to become more competitive over time, but this won’t threaten their businesses because, ultimately, a premium will continue to be maintained for certain workloads. Additionally, having the capacity to deliver that model for your workload is equally important. Believing that cutting-edge models are worth no premium is dishonest. How large that premium is remains to be seen.

If open-source models were banned or made illegal, labs would win big at the expense of innovation.

Reasoning has been proven to be profitable, based on service providers' profit margins being between 50% - 70%. Even if people leave hosted providers, that demand must flow toward their own purchased hardware. Given that reasoning is the dominant workload, demand far exceeds supply.

Now, how do our large labs perform in this world? I think the answer might be okay, but profit margins may not be so high.

Thinking they will fail and collapse is flawed. While I would love to give a perspective backed by more data, we do not have clear data on their specific profit margins; however, it is reasonable to speculate that their optimization capabilities for reasoning services have reached industry standards. Moreover, those companies with tens of millions or even hundreds of millions of monthly active users are not blatant Ponzi schemes or fraudulent projects, and their revenues continue to grow rapidly.

Image: Competitive landscape between large labs and open-source models

Return on Investment of Lab Models

While labs may still be unable to generate significant returns on SOTA models, this means the market won't reward newer models with better capabilities, nor will it be willing to pay for that premium. Considering that cutting-edge models represent the next class of problems AI can solve, betting against cutting-edge seems unwise. The cost of exiting this race will make it harder to catch up later (unless through distillation).

No matter how the SOTA race goes, reasoning still requires hardware, and the infinite demand for reasoning still must be met by someone’s hardware.

The Complexity of AI Trading

AI trading may be the most interesting trade in the current market because it has three characteristics:

  • Overlooked
  • Overcrowded
  • Severely misunderstood

Many market participants (including large institutional investors) fail to comprehend the complexities of technology and its impact on the bottom line of the balance sheet. Simplified headlines like Kimi K3 led to massive sell-offs of memory stocks, despite K3 being one of the largest memory-intensive models you can run.

These stock prices have risen, but they are trading at single-digit forward price-to-earnings ratios, with the market expecting their profits and demand to plummet in 2028/2029.

The New Economy of Infinite Demand

From a rationalism perspective, the existence of this infinite demand sounds crazy as a new economic commodity. Traditional economics suggests that if demand surges, supply will catch up to meet demand. However, considering the homogenized inputs (electricity, chips, land) and the uncertain outcomes (intelligence), we will never be satisfied with the existing computing power.

A world with infinite demand for computing power means we are entering a whole new realm.

Many computing power forecasts also do not account for the rise of robots over the next five years. Shorting computing power is equivalent to shorting robots. With the global population declining, our current economic model about more people becoming more productive is no longer valid (especially as many people's attention spans become shorter). Without AI and robots, there is no path to economic growth and creating more economic value.

Without AI, we will ultimately slow the progress of humanity. As a civilization, to reach this stage, we need to bring more computing power online. North America's capacity to bring computing power online has basically reached its limit. Canada and Australia are next. We will build data centers around the world until space runs out and then turn to space to build more data centers.

Asymmetrical Pricing of Supply and Demand

What is even more fascinating is that the supply is precisely measured as it is known through the foundry schedules in earnings calls, and the pricing is perfect (all supply will arrive on an exact timetable, without delay). In contrast, the demand side is measured in real-time and is underestimated. This creates a situation where the market believes in the ceiling of supply while significantly underestimating the demand situation (the underestimation is substantial, and today there are credible data).

Whenever you try to reason through AI trading, you need to ask yourself whether you believe the demand for computing power is limited or infinite. The answer to this question will guide your subsequent decisions.

Computing power is not a normal commodity. It is both a substitute for labor and a necessity for national security, and it is also a recursively self-replicating means of production. Believing that demand is limited is tantamount to believing that humanity has reached satisfaction; believing that demand is infinite implies that we are at the beginning of an economic paradigm shift.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink