
After Semiconductor Stocks Plummet 40%, Revisiting the Fundamentals of Compute Demand
TechFlow Selected TechFlow Selected

After Semiconductor Stocks Plummet 40%, Revisiting the Fundamentals of Compute Demand
When the compute order backlog of hyperscale cloud providers rises from $500 billion to $2 trillion, claiming this is a bubble requires stronger evidence.
Author: Kerman Kohli
Compiled by: TechFlow
TechFlow Editor's Note: After semiconductor stocks plummeted by 30-40%, many declared the AI bubble burst. But this judgment is built on a fatal assumption: that demand for compute is limited. From government military and scientific research to enterprise products and personal applications, all groups are competing for compute, and each group has a different price ceiling they are willing to pay. More critically, AI has a characteristic other infrastructure lacks—recursive demand: compute itself generates more demand for compute. When hyperscalers' compute order backlog grew from $500 billion to $2 trillion, claiming this is a bubble requires stronger evidence.

Core Question: Infinite Demand or Limited Demand
Depending on when this article is published, semiconductor and other AI/momentum-related stocks have fallen 30-40% from historical highs.
Many are willing to call this the top for semiconductors/AI/memory and celebrate victory for not participating.
They are likely celebrating too early.
In my view, the semiconductor/AI investment thesis boils down to this one question:
"Do you believe the demand for compute is limited or infinite?"
In conversations, I see too many people stuck in their localized experience of enterprise usage/adoption and generalize it to the broader market. I think the argument that enterprise adoption needs more time may indeed hold true.
However, this also creates a counter-incentive for smaller, more AI-native companies to defeat them with fewer employees, because AI, if used properly, can be much cheaper than expanding with manpower.
In any case, let's zoom out and stop viewing AI capital expenditure merely as enterprise demand. At a high level, AI construction is about bringing compute online.
Compute in turn can and will be used by the following groups:
- Governments for military and defense purposes
- Scientists for medicine and other frontier research
- Enterprises building new products and expanding without labor constraints
- Individuals empowered to do more (programming, design, creation, questioning)
These use cases and buyers are each willing to pay different prices for compute. While some may drop out due to compute prices, it is unwise to believe no one else has a higher budget. For some, compute expenditure is non-negotiable because this is a brutal race. Examples include sovereign governments and hyperscalers. For others, compute is a substitute for labor costs and still much cheaper (no legal overhead, management time, etc.).
Believing we are "overbuilding" or "building beyond capacity" means any of the above groups have reached a final state and are satisfied with the status quo. Simply put, it believes:
- Governments believe they do not need smarter weapons and defense capabilities
- Scientists are satisfied with the amount of research completed
- Enterprises believe they have done enough work on their product lines and do not want to grow further
- Individuals have reached a final state of curiosity and do not want to do more
If you truly believe any of the above, then you are right to say AI is a bubble and there will be overinvestment.
Each group is willing/able to pay different prices for compute, but the market will organize around demand to ensure quality/price can be met. Thinking intelligence is too expensive for all groups is a misconception, because those who can profitably orchestrate intelligence will continue to drive demand for it.
However, if you believe humans and the above groups will never be satisfied, then you must believe the demand for compute is infinite. We are currently in a great race around compute, and few people realize this.
Follow-up Question: What is the Return on Investment
After spending more time in the market, this is the biggest ongoing concern for investors regarding compute construction. Is the trend of hyperscalers spending all their free cash flow excessive, and are they betting their future?
Beyond hyperscaler capital expenditure, people are also very concerned about what revenue and profit large labs are generating? New open-source models threaten the value labs can extract from frontier models, making the situation more complex.
I will spend time discussing each one, but let's start with hyperscaler capital expenditure.
To those who think hyperscalers are miscalculating, I need you to understand this is far from the truth: public cloud services are ridiculously expensive, and they know how to squeeze every penny out of you. They have convinced a whole generation of companies to believe they cannot scale without them.
In return, they mark up the cost of normal non-CPU compute by 10-20 times. Additionally, you have to pay for logs, data transfer (egress), and 5 other services to complete basic work.
This game works because they trap you in their system. Bandwidth within the GCP/AWS kingdom is cheap, but once you transfer outward it rises significantly. For many hyperscalers, they need customers to stay in their ecosystem, otherwise there is a risk of losing business. Not having enough compute is fatal to their survival. When your customer data and compute are both with you, saying your GPUs are used up is completely unacceptable and would force them to slowly switch to your competitors. Hyperscalers have created an interesting dynamic where they can force customers to pay any price they want, and customers are powerless against it. Their tenants are so disconnected from bare metal, with huge lock-in effects, making migration a multi-year effort (if possible at all).
Here is a simple example to illustrate how crazy they are in this regard. I wrote a few months ago about how I built this $15,000 machine:
Building an AI Inference Machine

Figure: Author's self-built AI inference machine Source: Kerman Kohli / Substack
You can find the same GPU, lower configuration machine on GCP:
https://cloud.google.com/products/compute/pricing/accelerator-optimized
On-demand cost is about $3,248 per month, 3-year reserved is $1,444 per month.

Figure: Google Cloud GPU instance pricing example. Source: Google Cloud
My machine only has 128GB DDR5, but Google Cloud's has 180GB of some memory (they don't tell you if it's DDR4 or DDR5 haha).
Quick calculation:
- At on-demand pricing my machine will pay for itself within 4.6 months
- Three-year commitment payback period is 10 months
The math for other machines is similar. H200 clusters (GPUs released late 2024) will pay for themselves in less than 2 years. Wherever you look, you see very similar math. Of course this does not consider: land costs, ongoing electricity bills, financing costs and on-site staff, but it should serve as an illustrative example of how hyperscalers know how to price with huge premiums. Of course spot prices and committed prices differ again.
What makes this math even crazier is that GPUs from 5 years ago are: a) holding value b) leasing costs are rising!
Hardware is very likely not to depreciate, but to appreciate from here on. Although new chips with higher compute efficiency are launched, the efficiency they lack will be compensated by rising memory costs.
I conceptualize hardware as a two-component game, where one component (compute) technically becomes less valuable, but is offset by another component (memory), which becomes more valuable over time.
If this is the case, their capital expenditure ROI is higher than anyone remotely expected. Regarding hyperscaler credit risk, this tweet from Gavin Baker summarizes it well:

Figure: Gavin Baker tweet image. Source: X / @GavinSBaker
Now you might say, how do we know this demand is sufficient? I mean I cannot model every scenario for every customer, but at some point you need to have the humility to say, people lining up to pay is the strongest signal, you believe they are rational actors spending on positive ROI efforts.
If we look at it from this angle, the order backlog grew from $500 billion at the beginning of 2025 to far beyond $2 trillion in just 1.5 years. When this is customer-driven, saying all of this is fake/not positive ROI becomes far-fetched. Now, the counterargument is that labs account for a large part of this, but this view is wrong. According to EpochAI, frontier labs account for a part, but not all of global compute demand.

Figure: Compute order backlog grew from $500 billion to $2 trillion. Source: EpochAI

Figure: Global compute demand composition. Source: EpochAI
No matter what you believe, the fact of a $2 trillion backlog should indicate something. Thinking trillions in expenditure does not reflect a structural shift but rather excessive bubble is an interesting viewpoint.
Many investors like to reason by analogy, comparing it to internet construction, railways, or past infrastructure projects. I understand the logic here, but it ignores a key feature of AI: recursive demand.
For railways or the internet, you need more people to adopt the technology, and then ensure everyone's upper limit of using the technology, to ensure sufficient diffusion in the economy. AI does not have these dynamics. In this race, compute can generate its own demand for compute, and an individual or organization's compute limit is actually infinite. If you find a useful, positive ROI use case, you can continue to invest in compute, and it will become a money-making machine.
Where this gets very tricky is that different people have very different experiences using AI. Most people in the world use it as a single prompt Q&A machine. For people like me, as agent engineering becomes more capable, they are becoming indispensable, letting me do more things.
My compute expenditure continues to rise, and will continue to rise, because I have discovered more positive ROI use cases. No matter how much revenue AI generates, the cost savings it creates are undeniable, which drives the case for most end customers.
Uncertainties: OpenAI / Anthropic
Continuing from the last point, we can see the demand backlog is crazy. But how real is the demand from labs (a significant portion of compute demand). Now this is where I think the answer is less clear but still deducible. I want to break this answer down into inference and training.
If we consider a new SOTA (State-of-the-Art) model costs up to hundreds of millions of dollars, then we can say it is an investment asset, generating some useful lifecycle value over time through inference (despite a steep depreciation curve).
As a counterforce, you will have open-source models diffusing into the market and competing with frontier labs for compute at cheaper costs. Whether these open-source models are distilled or not is not important for understanding the dynamics.
So the dynamic we must question is what happens when SOTA models come out? The reality is not everyone will always use them to solve every problem. However, given they have SOTA capabilities, they can solve problems current model classes cannot, and you are willing to pay a premium for this.
You could say models like Kimi K3 change this dynamic because they are open-source, but people forget an important fact: SOTA models are very large, and the hardware required to run them far exceeds what any home model can do. Kimi K3 itself requires close to 1.5TB - 2TB of memory. Good luck finding it.
What makes the model more interesting is that Kimi exhausted capacity after opening the gates to K3. Of course a model exists out there, but someone still needs to serve it. It still must run on capable hardware. Labs will indeed be forced to become more competitive over time, but this will not threaten their business, because eventually the premium for certain workloads will remain. Furthermore, being able to provide capacity for that model for your workload is equally important. Thinking frontier models are not worth any premium is dishonest. How large this premium is remains to be seen.
If open-source models are banned or illegal, then labs will win big at the cost of innovation.
Inference has proven profitable, with service provider margins between 50% - 70%. Even if people leave hosted providers, this demand must flow to them buying their own hardware. Given inference is the dominant workload, demand far exceeds supply.
Now regarding our large labs, how do they perform in this world? I think the answer might be okay, but margins might not be that high.
Thinking they will fail and fall is wrong. Although I would like to give a view supported by more data, we do not have clear data on their specific profit margins; however it is reasonable to speculate that their inference service optimization capabilities have reached industry standards. Furthermore, those companies with tens of millions to hundreds of millions of monthly active users are not naked Ponzi schemes or fraud projects, and their revenue is still growing rapidly.

Figure: Competitive landscape of large labs and open-source models
ROI of Lab Models
Labs may still be unable to generate significant returns on SOTA models, but this means the market will not reward new models with stronger capabilities, nor is it willing to pay for this premium. Considering frontier models represent the next class of problems AI can solve, betting against frontier seems not a good idea. The cost of exiting this race is harder to catch up later (unless through distillation).
Regardless of the SOTA race, inference still requires hardware, and the infinite demand for inference must still be met by someone's hardware.
Complexity of AI Trading
AI trading might be the most interesting trade in the current market, because it simultaneously has three characteristics:
- Overlooked
- Overcrowded
- Severely misunderstood
Many market participants (including large institutional investors) cannot understand the complexity of the technology and its impact on the balance sheet bottom line. Simplified headlines like Kimi K3 led to massive sell-offs in memory stocks, even though K3 is one of the largest memory-intensive models you can run.
These stock prices have risen, but they trade at single-digit forward P/E ratios, with market expectations that their profits and demand will plummet in 2028/2029.
New Economy of Infinite Demand
From a rationalist perspective, saying there exists this new economic commodity with infinite demand sounds crazy. Traditional economics believes that if demand surges, supply will catch up to meet demand. But considering homogeneous inputs (electricity, chips, land) produce uncertain outcomes (intelligence), we will never be satisfied with existing compute.
A world of infinite compute demand means we are entering a completely new world.
Many compute forecasts also do not consider the rise of robots in the next five years. Shorting compute is shorting robots. As global population declines, our current economic model about more people becoming more productive no longer holds (especially as many people's attention spans are getting shorter). Without AI and robots, the economy has no path to growth and creating more economic value.
Without AI, we will ultimately slow down human progress. As a civilization to reach this stage, we need to bring more compute online. North America's compute online capacity has basically reached its limit. Canada and Australia are next. We will build data centers all over the world until space runs out, then turn to space to build more data centers.
Asymmetric Pricing of Supply and Demand
What is more fascinating is that supply is precisely measured, because it is known through foundry schedules in earnings calls, and priced perfectly (all supply will arrive on exact schedule, no delays). While the demand side is measured in real-time, and underestimated. This creates a situation: the market believes in the supply cap, while underestimating the demand situation (underestimation is large, and today there is real credible data).
Whenever you try to reason about AI trading, you need to ask yourself whether you believe compute demand is limited or infinite. The answer to this question will guide your subsequent decisions.
Compute is not an ordinary commodity. It is both a substitute for labor, a necessity for national security, and moreover recursive self-reproducing means of production. Believing demand is limited equals believing humans are already satisfied; believing demand is infinite means we are standing at the starting point of an economic paradigm shift.
Join TechFlow official community to stay tuned
Telegram:https://t.me/TechFlowDaily
X (Twitter):https://x.com/TechFlowPost
X (Twitter) EN:https://x.com/BlockFlow_News














