NextFin

Open Models Are Eating AI's Model Layer, but the Hyperscalers Still Own the Tollbooth

Summarized by NextFin AI
  • Open-weight AI models from Alibaba, Meta and Chinese labs now match frontier systems at a fraction of the cost, compressing the model-layer value while hyperscalers keep collecting infrastructure tolls.
  • Four hyperscalers plan $600B-$725B capex in 2026, up 60%-77% from 2025, betting AI inference demand compounds fast enough to outrun GPU depreciation.
  • Model commoditization evidence is hard: Qwen logged 399.4M downloads (18.5% of top-1,000 traffic), while DeepSeek V3 trained for ~$5.6M versus $100M+ for GPT-4-class runs.
  • The real risk is the capex cycle, not commoditization: overbuilding could turn cash-generative franchises into capital-intensive utilities before revenue catches up.

NextFin News - Open-weight AI models from Alibaba, Meta and Chinese labs now match frontier systems at a fraction of the cost, and the question haunting the $600 billion-plus AI buildout is whether the cloud giants' moat is about to evaporate. The short answer: the threat is real, but it is aimed at the wrong target. Open models compress the value of the model layer itself — the OpenAIs and Anthropics of the world — while the hyperscalers that own the data centers keep collecting the toll, whether the token running on their GPUs costs $30 or $1 per million.

The uncomfortable question for investors is not whether intelligence is becoming a commodity. It is whether the tollbooth can generate enough traffic to pay for a bridge that costs three-quarters of a trillion dollars a year.

The Numbers That Define the Bet

The stakes are set by the capital expenditure plans of the four hyperscalers. Amazon, Alphabet, Microsoft and Meta plan to spend roughly $600 billion to $725 billion on capital expenditures in 2026, a 60 percent to 77 percent jump from the roughly $390 billion to $410 billion spent in 2025. Amazon alone guided to about $200 billion, up from $125 billion. Alphabet guided to $175 billion to $185 billion, up from $91 billion. Meta guided to $115 billion to $135 billion, up from $72 billion. Microsoft is expected to spend $110 billion to $120 billion, up from about $90 billion.

That spending is a bet that demand for AI inference will keep compounding fast enough to outrun the depreciation on GPUs that become obsolete within a few years. And into that bet steps the open-model wave.

The evidence of commoditization at the model layer is now hard. In a 2026 analysis of the 1,000 most-downloaded models on Hugging Face's model hub, Alibaba's Qwen family accounted for 399.4 million downloads across its top models — 18.5 percent of all top-1,000 traffic and nearly four times Google's footprint on the hub. Meta's Llama family, often treated in Western markets as the dominant open-weight family, logged 29.6 million downloads in the same basket, about one-thirteenth of Qwen's volume. Google followed at 104.2 million and OpenAI at 91.6 million. Seventy-one and a half percent of the top models ship under permissive open licenses, with Apache 2.0 alone covering 49 percent. Qwen-based derivatives number 151,448 on the hub, 2.6 times Meta's total footprint.

The cost asymmetry underneath those download figures is starker. DeepSeek's V3 base model was trained for an estimated $5.6 million in compute, against industry estimates of more than $100 million for GPT-4-class runs. Its R1 reasoning model reportedly used only $294,000 of GPU time. The International Institute for Strategic Studies calculated that V3's inference runs 12.5 times cheaper than Anthropic's Claude 3.5 Sonnet and more than 15 times cheaper than OpenAI's GPT-4o. OpenAI's chief executive said in early 2025 that DeepSeek's R1 runs 20 to 50 times cheaper than OpenAI's comparable model.

The price tags tell the same story. On OpenRouter, a routing platform widely used as a proxy for model market share, frontier models commanded $25 to $30 per million tokens as of late May 2026 while budget open models traded under $1 per million — a 25-fold gap that customers kept paying. Anthropic, on that platform, processed about 11 percent of tokens but captured 42 percent of estimated model spend. Companies, it turns out, do not pay for benchmark points. They pay for an agent that does not fall apart in production.

But the gap is narrowing, and the direction of travel is what matters. The share of calls on OpenRouter sent to models priced under $1 per million tokens rose from 18 percent in January 2026 to 41 percent in June. The share of tokens processed by Chinese open models such as DeepSeek, Qwen and Kimi went from under 2 percent in late 2024 to about 61 percent in May and June 2026; counting all open-weight models, the latest weekly snapshot sits near 69 percent. The Stanford HAI AI Index found that inference cost for GPT-3.5-level performance fell more than 280-fold in two years; a16z pegs the decline at roughly 10 times per year for any fixed capability level.

The Tollbooth Thesis: Why Hyperscalers Win the Commoditization War

The cleanest way to think about the AI stack is a highway. The AI labs manufacture the cars — the models that do the actual work, with real margins in the frontier tier. The hyperscalers own the tollbooth. Every car that crosses the bridge pays the toll, whether it is a Ferrari or a used Honda.

Token optimization — the shift from pointing every workload at the single best model toward routing cheap open models at routine tasks and escalating only hard requests to frontier systems — looks bearish for the labs and should compress the whole stack. In practice it is one of the most bullish structural setups for Microsoft, Amazon and Google.

The reason is arithmetic. When a company routes a workload to an open-weight model — GLM, DeepSeek, Qwen, Llama — running on a hyperscaler's managed inference, the model-provider margin collapses toward zero. Nobody charges a brand premium for an open weight. But the token still has to run on somebody's GPUs, inside somebody's data center, behind somebody's managed API with its security, compliance, logging and service-level guarantees. That somebody is the hyperscaler. And the infrastructure margin does not care whether the token came from a $50-per-million frontier model or a $1-per-million open-weight model.

The hard numbers back this up. AWS ran at roughly a 39.5 percent operating margin in the first quarter of 2026, on $37.59 billion of revenue. Google Cloud, which lost money for years, posted a 32.9 percent operating margin in the same quarter, up from 17.8 percent a year earlier, on revenue that grew 63 percent to $20.03 billion. Azure grew 40 percent in Microsoft's third fiscal quarter of 2026, its third straight quarter above 38 percent, taking FY2025 Azure revenue to $75 billion. Those margins are being printed while the model war rages underneath them.

Microsoft's own disclosures show the demand side compounding. It processed more than 100 trillion tokens in a single quarter in 2025, up five times year over year, with a record 50 trillion in one month. By its fiscal third-quarter 2026 call, more than 300 customers were on track to process more than a trillion tokens each on Azure AI Foundry that year, accelerating 30 percent quarter over quarter. Google went from 480 trillion tokens a month at its I/O developer conference in May 2025 to 980 trillion by July to 1.3 quadrillion by October, and disclosed in its first-quarter 2026 filing that its first-party models alone were processing more than 16 billion tokens per minute via direct API, up 60 percent in a single quarter.

"People are really saying, you know, it's kind of a meme now, but 'My company spent my entire 2026 budget in Q1. Can you make this more efficient?'" OpenAI chief executive Sam Altman said on stage at the company's Intelligence at Work event, adding that the complaint went from an issue that never came up at the beginning of the year to "all of a sudden, a huge issue."

That is not a demand problem. That is a budgeting problem created by demand arriving faster than finance departments planned for. The tollbooth collects on every extra mile.

The Second-Order Effect: Openness Accelerates the Infrastructure Boom

The counter-intuitive read is that open models do not shrink the hyperscalers' opportunity — they expand it. JPMorgan Asset Management's January 2026 outlook documented GPU rental rates falling 20 percent to 26 percent across the prior year while compute demand kept rising, a combination that points the same way: as model rents compress, demand for chips, cloud and power rises faster. Openness makes intelligence less scarce, and cheap intelligence is used in far greater volume.

This is the second-order transmission that the market underweights. The first-order effect of an open model is obvious: the price of a unit of intelligence falls. The second-order effect is that demand for units of intelligence is elastic — when inference gets cheap enough, companies stop rationing it. They run the agent in a loop. They let it read the whole codebase. They re-run it five times and vote on the answer. A single coding-agent session now chews through millions of tokens of context where a chatbot query used a few thousand.

The consequence is that the hyperscalers' revenue base broadens even as the model layer thins. Enterprises that could never justify a frontier API bill for routine classification, extraction, summarization or support tickets can run those workloads on cheap open models hosted in the same cloud. The marginal workload that never existed at $30 per million tokens exists at $1. The hyperscaler bills for the compute either way.

Scenario analysis published on the economics of the switch puts the break-even point starkly: if companies re-spend about a third of what they save by moving to cheaper models back into additional token consumption, the hyperscaler earns as much profit as before the switch; above that reuse rate, the infrastructure provider earns more. Reports of annual AI budgets exhausted in a single quarter, and token usage growing five to seven times a year, suggest reality sits to the right of that break-even bar. In that zone, the premium that used to go to the frontier developer disappears, and hyperscaler profit and server operations fill the space.

The Real Risk Is Not Commoditization — It Is the Capex Cycle

Which brings the analysis to the actual threat. The open-model wave is not the thing that breaks the hyperscalers. The thing that breaks them, if anything breaks them, is the capital intensity of the response.

Goldman Sachs, in research notes published in late April and early May 2026, warned that Microsoft, Amazon, Alphabet and Meta may collectively spend more than $600 billion on capital expenditures, potentially consuming 100 percent or more of their operating cash flow. That is the vulnerability: not that open models erode pricing power, but that the buildout to host them — and to stay in the frontier race — turns these cash-generative franchises into capital-intensive utilities before the revenue catches up.

The assets are short-lived. Microsoft's finance chief said roughly two-thirds of its second-quarter fiscal 2026 capital expenditure, which reached $37.5 billion, went to short-lived assets, primarily GPUs and CPUs, because customer demand exceeds supply. A GPU depreciates over three to six years depending on the company's accounting; a frontier model can be outclassed in eighteen months. If utilization slips — if the inference demand does not compound at the pace the capacity assumes — the depreciation charge hits margins long before any pricing power does.

This is a cyclical risk layered on top of a structural shift. Structurally, value is migrating down the stack from the model layer to the infrastructure layer, and that migration favors the hyperscalers. Cyclically, the industry is in a capacity race where the winner is whoever built just enough, and the loser is whoever built too much. History's lesson from every infrastructure supercycle is that overbuilding, not commoditization, is what destroys returns.

The Counter-Thesis: What If the Tollbooth Gets Bypassed?

The strongest case against the tollbooth thesis deserves to be stated fully. It runs like this: open weights are, by definition, portable. A company that runs Qwen or Llama on its own GPUs, or on a neocloud such as CoreWeave or Oracle, has no structural reason to route the workload through Azure, AWS or Google Cloud. The more standardized and open the model layer becomes, the more the workload becomes a pure commodity that migrates to the cheapest available GPU, wherever it sits. The hyperscaler's moat was never concrete and power contracts — it was the integration of proprietary models, proprietary tooling and proprietary distribution. Remove the proprietary model from the center of the stack, and the rest becomes contestable.

There is evidence for this view. Meta retired its hosted Llama API on July 6, 2026, a signal that even the company that did the most to popularize open weights concluded there was no business in giving the product away through its own hosted service, pointing developers instead to third-party hosts. Oracle is building multi-hundred-thousand-GPU clusters explicitly to court the frontier labs. Neoclouds are pre-leasing capacity directly to OpenAI and others, bypassing the big three. And Morgan Stanley has argued that open-source models threaten to commoditize the underlying technology layer, making it harder to maintain a durable advantage based on the models themselves.

The answer to the counter-thesis is that portability is real but incomplete. Running an open model in production is not the same as downloading weights. Production requires security, compliance, logging, audit trails, identity management, data residency, uptime guarantees and integration with the rest of the enterprise stack — the unglamorous plumbing that hyperscalers already own and that enterprises already trust with their most sensitive workloads. The model may be portable; the governance layer is not. That is why the routed-token economy still converges on the big clouds even when the model is free.

But the counter-thesis identifies the genuine falsifying signal, and it should be stated as a concrete threshold rather than a vague watch item. If neocloud and on-premises open-model deployments account for more than 20 percent of enterprise AI inference workloads within the next four quarters — up from the low single digits today — the integration moat is eroding faster than demand growth can compensate, and the tollbooth thesis is wrong. Portability would be beating integration, and the hyperscalers' margin would be contestable after all.

Who Wins, Who Loses, and What to Watch

Cash in the mechanism. If the analysis holds, the winners and losers of the open-model wave are not the ones the headline suggests.

The exposed: the frontier model labs whose pricing power rests on a performance gap that keeps closing. A 25-fold price gap between frontier and open models is a moat only for as long as customers believe the extra performance is worth 25 times the money. For routine enterprise work, increasingly, it is not. The labs must either defend the gap with genuinely differentiated capability — agentic reasoning, proprietary data integration, trust and safety — or watch routine volume migrate to open weights.

The beneficiaries: the hyperscalers, but with a large caveat. They capture infrastructure rent on every routed token, open or closed, and open models expand the total volume of tokens that run in the cloud. Microsoft, Amazon and Google are best positioned because they combine scale, distribution and the governance layer enterprises require. Meta is the hybrid case — both a massive consumer of cloud capacity and the engine of open architecture through Llama, now pivoting toward proprietary models under its Superintelligence Labs.

The caveat is the capex cycle. The structural shift toward infrastructure is real, but the cyclical risk of overbuilding is equally real. Investors should treat the hyperscalers not as a single trade but as a question of execution: who can keep utilization high enough, long enough, to depreciate hundreds of billions of capacity without crushing returns.

The forward look splits by time horizon. In the short term, over the next two to three earnings seasons, watch utilization and the capex-to-revenue ratio to see whether token growth is keeping pace with capacity additions. In the medium term, watch the neocloud share of AI inference against the 20 percent threshold above, and the persistence of the frontier-open price gap. In the long term, watch whether value migrates further up the stack to the application and workflow layer, where proprietary context and distribution may ultimately matter more than either models or infrastructure.

Scenarios: the base case is that model rents compress while chip, cloud and power demand rise faster, and the hyperscalers absorb the open-model wave as volume growth. The upside case is that cheap inference unlocks demand nobody modeled, and the tollbooth collects on an order of magnitude more traffic than anyone expected. The downside case is that capacity comes online faster than demand, utilization slips, and the short-lived GPU assets depreciate into a margin hole that no amount of routing can fill.

The open-model threat to the hyperscalers is smaller than the panic suggests — but the capex threat they brought on themselves is larger. Intelligence may be becoming a commodity. The bill for the data centers is not.

Explore more exclusive insights at nextfin.ai.

Search
NextFinNextFin
NextFin.Al
No Noise, only Signal.
Open App