NextFin

Cheap Tokens, Costly Chips and the Missing AI Payoff

Summarized by NextFin AI
  • AI token prices collapsed ~95% in two years, from $30 per million input tokens in March 2023 to under $0.50 in 2026, while hyperscaler capex is projected to rise 62% to $630 billion in 2026.
  • Four hyperscalers drive the spending surge: Amazon $200B, Alphabet $175B-$185B, Meta $115B-$135B, Microsoft $110B-$120B, with UBS estimating they will spend 102% of cloud revenue on AI infrastructure.
  • Pure-play AI vendors remain unprofitable: OpenAI's $40B run rate equals only ~6% of hyperscaler capex, with projected 2026 losses of ~$14B; Anthropic's revenue grew to >$47B but is still racing toward profitability.
  • June 2026 chip selloff erased >$1.3 trillion from the sector: Philadelphia Semiconductor Index fell 10.3%, Nvidia shed ~$330B, while S&P 500 fell 2.6% and Nasdaq 100 dropped 4.8%.

NextFin News - Artificial intelligence has never been cheaper to run, and that is the problem. The price of producing GPT-4-class output has collapsed from about $30 per million input tokens in March 2023 to under $0.50 per million in 2026 - a decline of roughly 95% in two years and about 1,000-fold over three. Yet the bill for the infrastructure behind that cheap intelligence keeps climbing: Amazon, Alphabet, Meta and Microsoft plan to spend as much as $630 billion on capital expenditures in 2026, about 62% more than the record $388 billion laid out in 2025. The widening gap between collapsing token prices and soaring chip bills is the central tension of the AI buildout, and it is why the payoff investors were promised keeps moving further into the future.

The Two Curves That Do Not Meet

The AI economy is being pulled apart by two curves moving in opposite directions. On one side, the unit cost of intelligence is falling faster than almost any input in the history of computing. When GPT-4 launched in March 2023, it cost $30 per million input tokens and $60 per million output tokens. By mid-2026, open-weight models matching that quality bar were available for roughly $0.10 per million input tokens and $0.40 per million output tokens - a drop of about 99% on both measures once the full three-year arc is counted. Nvidia's Blackwell architecture is claimed to deliver up to 25 times lower cost and energy per inference token versus the prior Hopper generation for large mixture-of-experts models, before quantization and better serving software are even factored in.

On the other side sits the heaviest capital-spending cycle in technology history. The four largest hyperscalers guided to as much as $630 billion in combined capital expenditures for 2026: Amazon at $200 billion, up from $125 billion; Alphabet at $175 billion to $185 billion, up from $91 billion; Meta at $115 billion to $135 billion, up from $72 billion; and Microsoft at $110 billion to $120 billion, up from $90 billion. A baseline model from Goldman Sachs puts annual AI capital expenditure at $765 billion in 2026, rising to $1.6 trillion in 2031, with roughly $7.6 trillion deployed across compute, data centers and power between 2026 and 2031. UBS estimates that Amazon, Alphabet and Microsoft will collectively spend about 102% of their cloud revenue on capital expenditures in 2026 - recycling essentially all of cloud income back into AI infrastructure.

Between those two curves lies the missing payoff. The pure-play AI vendors consuming this infrastructure are growing rapidly but from bases that remain small relative to the capital being deployed on their behalf. OpenAI's annualized revenue run rate reached $40 billion as of August 2026 - impressive for a company that barely had consumer products three years earlier, but equal to roughly 6% of the hyperscalers' projected 2026 capex. On a booked basis, leaked audited financials showed $13.07 billion of 2025 revenue against a $20.9 billion operating loss, and the company is projected to lose roughly $14 billion in 2026. Anthropic's annualized revenue reportedly grew from about $9 billion at the end of 2025 to more than $30 billion in April 2026 and more than $47 billion by mid-May 2026, yet it too is racing toward profitability rather than arriving there.

Why Cheap Tokens Do Not Translate Into Profits

The obvious bull argument is that falling prices are working exactly as intended: cheaper tokens expand the market, and total spending rises even as unit prices fall. This is the Jevons paradox, the 19th-century observation that as a resource becomes more efficient and cheaper, consumption rises more than enough to offset the savings. Microsoft chief executive Satya Nadella invoked it directly when a low-cost model shook the market: "Jevons Paradox strikes again." The data supports the demand side of the story. Industry research projected worldwide generative AI spending of $644 billion in 2025, a 76.4% increase from 2024, even as per-token prices fell roughly 1,000-fold. Agentic workloads consume approximately 1,000 times more tokens than standard chat interactions for equivalent tasks, because agents resend full context at every reasoning step, load tool schemas, retry failed steps and run around the clock.

But the Jevons defense contains its own trap. When the unit price of your product falls 95% while your infrastructure costs rise 62%, demand must grow by more than the price decline just to stand still - and then grow further still to cover the new capital being poured in. A company whose revenue is tied to token volume needs usage to roughly double just to offset a 50% price cut; at a 95% price decline, usage must expand roughly twentyfold before the revenue line recovers the ground that pricing gave up. The hyperscalers are not merely holding existing ground. They are adding hundreds of billions in new capacity that must be filled, depreciated and returned to investors at a cost of capital that has risen sharply alongside long-dated Treasury yields.

This is where the second-order effect bites. Cheap tokens do not only pressure the AI labs that sell them. They transmit deflation through the entire stack. Every custom-silicon team, every cloud provider and every application vendor now prices against a commodity floor set by open-weight models hosted near cost. Nvidia can still charge a premium for the frontier - top-end graphics processing units and high-bandwidth memory remain sold out through 2026, with no real relief arriving until 2028 - but the mix of demand is shifting away from top-end training parts toward inference-optimized silicon, where margins are thinner and competition is fiercer. The revenue that justified buying chips at 2023 prices must now be earned at 2026 prices, with the same or heavier fixed costs.

"During the training phase, the cost of AI infrastructure and token generation is extraordinarily high, but in the current inference stage, the economics are significantly better," said David Miller, a senior portfolio manager at Catalyst Funds. "The net use of AI delivers a positive return on investment for companies, at least over the long term."

The long term is doing a lot of work in that sentence. In the short term, the mismatch is visible in the market. The sequence began on June 3, 2026, when Broadcom reported second-quarter AI semiconductor revenue of $10.8 billion, up 143% year over year, yet left its full-year AI chip revenue target unchanged at $56 billion; shares fell more than 13% in extended trading after the company missed Wall Street's second-quarter revenue estimate and declined to raise its 2027 forecast. Two days later the unwind broadened into one of the most concentrated selloffs in recent market history. More than $1.3 trillion in market value was erased from the global chip sector. The Philadelphia Semiconductor Index sank 10.3%, its deepest one-day loss since March 2020. Nvidia alone shed roughly $330 billion in market value, while Taiwan Semiconductor, Broadcom and Micron each lost more than $100 billion. The iShares MSCI South Korea ETF, a live read on memory and AI supply-chain stress, fell 14.1%, its worst day since March 2020. The S&P 500 fell 2.6% and the Nasdaq 100 dropped 4.8%, but the damage was concentrated: this was not a broad-market collapse, it was a repricing of the market's most crowded trade.

Cyclical Panic, Structural Deflation

It matters to separate what is cyclical from what is structural, because the two point in opposite directions. The capex cycle is cyclical. Semiconductor demand has always oscillated between shortage and glut, and the June 2026 wipeout fits the classic pattern of a crowded trade meeting a single disappointing data point. When financing conditions tighten - the 30-year Treasury yield topped 5.3% in mid-August 2026, its highest level since 2007 - the discount rate applied to profits expected a decade out rises, and valuations built on distant payoffs compress first. This leg will mean-revert: if usage keeps compounding and a few killer applications emerge, the chips will be filled, utilization will rise, and the panic will look excessive in retrospect. Three historical-cycle comparisons anchor the call. The 2018 memory downturn saw prices fall more than 50% before rebounding on supply discipline. The 2020 pandemic shortage swung to oversupply by 2023 as capacity caught up. The 2022 crypto crash erased mining demand almost entirely, yet semiconductor revenues recovered as AI training demand emerged in 2023. In each case, the cycle reverted because the underlying demand driver - PCs, smartphones, cloud - was intact.

Token-price deflation is different. It is structural, and it will not self-correct. Three forces compound on top of one another, and none reverses on its own. Algorithmic efficiency - quantization, speculative decoding and mixture-of-experts routing - cuts compute per token five to twentyfold without quality loss. Cheaper, faster hardware raises inference throughput per dollar an estimated 25-fold or more over three years. Open-weight competition from Llama, DeepSeek, Qwen and Mistral sets a near-zero-margin price floor that closed labs must price against. This is a regime change in the economics of intelligence, not a temporary shortage. The correct analogy is not the memory cycle; it is the cost curve of computing itself. Moore's Law did not reverse when chip prices fell; it accelerated. Similarly, token prices are not going back to $30. Any business model that requires them to is not waiting out a cycle - it is betting against the direction of technological progress.

The counter-thesis is serious and deserves its due. The bear case, articulated by investors including Michael Burry, who warned that Nvidia's credit default swaps had gone "parabolic," is that the entire buildout rests on circular financing: chipmakers effectively funding their own customers through financing guarantees tied to data-center projects, with revenue that never reaches an end user. Nvidia was reported to be in talks to provide a roughly $250 billion financial backstop for OpenAI's 10-gigawatt data-center campus in southern Ohio, and Wall Street firms were said to be working on a plan to mobilize up to $500 billion in third-party capital for AI infrastructure. Those discussions have since been scaled back - by mid-August the Ohio guarantee had been reduced to under $120 billion, focused initially on the project's first phase - but the structure of the concern remains. OpenAI's projected cumulative losses of roughly $115 billion through 2029, before a swing to cash-flow profitability around 2029-2030 at more than $125 billion in annual revenue, illustrate the scale of the gap. For context, Amazon burned roughly $3 billion cumulatively in its first decade and Uber about $25 billion before reaching GAAP profitability; OpenAI's projected burn is an order of magnitude larger, funded at an $852 billion valuation.

The strongest answer to the bear case is that the financing is not as circular as it looks, and that the demand is real even if the monetization lags. Hyperscalers are not start-ups living on promised revenue. They are cash-generating enterprises with diversified businesses - search, cloud, advertising, e-commerce - that can absorb years of heavy investment the way Amazon absorbed years of logistics buildout before retail economics turned. Alphabet reported accelerating cloud growth even as its capex forecast rose, and the hyperscalers' executives have repeatedly expressed confidence that the bets will pay off. The Jevons dynamic is not theoretical: total token consumption is rising far faster than prices are falling, which means real work is being shifted onto AI infrastructure. Agents, long-context analysis and automated enterprise workflows are still in their first inning, concentrated in a few enterprises and heavy users rather than diffused across the economy.

But that answer carries a condition, and it is the condition the market is now testing. The bears are right if token prices keep falling faster than usage grows for long enough to exhaust the hyperscalers' financing capacity before the applications arrive. The bulls are right if usage compounding outruns price deflation and at least a few applications generate returns that cover the cost of the capital deployed. The falsifying signal is concrete: if the hyperscalers' combined AI-related revenue - cloud AI services plus AI-driven advertising and productivity revenue - does not reach at least 40% of combined AI capital expenditure within two years, the payoff thesis breaks and capex guidance will be cut. A simpler, more immediate tell: if any of the four hyperscalers reduces its forward capex guidance, the circular-financing fear becomes self-fulfilling, because the first cut is the signal that the market has been waiting for.

What Comes Next: Three Time Horizons

In the short term - the next six to twelve months - expect volatility, not collapse. The June 2026 wipeout demonstrated how quickly a crowded trade can unwind, but it also showed that the broader market absorbed the shock: the S&P 500's 2.6% decline against the semiconductor index's 10.3% drop is the signature of a sector rotation, not a systemic break. Top-end GPUs and high-bandwidth memory remain sold out through 2026, which means the chipmakers' near-term order books are still full. The risk is not a demand cliff; it is a multiple compression that can take out 20% to 30% of valuations before fundamentals catch up.

In the medium term - two to four years - the winners will be determined by who converts infrastructure into recurring revenue, not by who spent the most. The companies that embed AI into workflows with high switching costs, that price on value rather than tokens, and that own the customer relationship will capture the surplus that token-price deflation strips from the infrastructure layer. The exposed parties are the pure-play model vendors burning cash against a falling price floor, and the memory and equipment suppliers whose revenues are most leveraged to a single capex cycle. UBS projects hyperscaler capital expenditure reaching $1.009 trillion in 2026, $1.447 trillion in 2027 and $1.619 trillion in 2028 - three times more in three years than in the previous six combined. That spending cannot all be right. Some of it will be stranded, and the market will sort out which projects before the decade ends.

In the long term - five years and beyond - the structural call dominates. Intelligence will keep getting cheaper, and the economy will adapt to cheap intelligence in ways that are hard to predict from here. The beneficiaries will be the users of intelligence, not necessarily its producers - the same way the beneficiaries of falling computing costs were the companies that applied computing, not the hardware vendors of each cycle. If the Jevons dynamic holds, total spending on AI will keep rising even as the price per unit of intelligence approaches zero, and the capex that looks excessive today will look like the price of admission to the next computing platform. If it does not hold, the write-downs will be historic.

Three signals will tell the story before the earnings do. First, the hyperscaler capex guidance for the coming year - any cut is the canary. Second, the ratio of AI revenue to AI capital expenditure at the largest cloud providers - it must rise, and it must rise fast. Third, the token-price floor: if open-weight hosting keeps pushing the frontier toward $0.01 per million tokens, as some projections suggest, then every revenue model denominated in tokens must be rebuilt from scratch. The market has priced a version of the bull case in which cheap tokens create infinite demand. The alternative - cheap tokens creating infinite competition - is the risk that has not yet been fully priced.

The AI buildout is not a bubble in the simple sense of prices detached from any reality. It is something more subtle: a real technological revolution arriving with a pricing problem. The chips are costly, the tokens are cheap, and the payoff is missing because the revenue that was supposed to justify the spending is being competed away by the very efficiency the spending bought. Intelligence is becoming a commodity faster than the companies selling it can become profitable.

Explore more exclusive insights at nextfin.ai.

Search
NextFinNextFin
NextFin.Al
No Noise, only Signal.
Open App