NextFin

Sandra Rivera Bets the AI Capex Boom's Next Wave Belongs to Inference Chips

Summarized by NextFin AI
  • Sandra Rivera, chair of VSORA and former Intel DCAI head, frames the ~$732.5 billion 2026 AI infrastructure buildout as a transition point, with value migrating from training compute to inference efficiency rather than a peak to fear.
  • Hyperscaler capex hit $301 billion in H1 2026, with Microsoft, Meta, Amazon and Alphabet guiding to $732.5 billion for the full year and $934.5 billion consensus for 2027, approaching a $1 trillion annual run rate.
  • Gartner forecasts the semiconductor market to grow 92% in 2026 to $1.56 trillion, with memory revenue more than tripling to $837 billion and AI data centers rising from 36.5% of industry revenue to over 53% by 2030.
  • The thesis favors specialized inference accelerators and the infrastructure layer (memory, packaging, power, interconnect) over general-purpose GPUs for steady-state workloads, though a capex slowdown remains the key falsifying risk.

NextFin News - Sandra Rivera, chair of Paris-based AI chip startup VSORA and the former head of Intel's Data Center and AI Group, sees the roughly $732.5 billion AI infrastructure buildout of 2026 not as a peak to fear but as a transition point. The spending is real — hyperscalers have committed 39% more than analysts expected just months ago — but in Rivera's reading, the value is migrating from training compute to inference efficiency. The next leg of the semiconductor cycle, she argues, will reward architectures built for one job: running trained models cheaply, quickly, and at scale.

The Spending Wave and the Inference Pivot

Rivera is not an outside commentator on this market. She spent 23 years at Intel, rising to executive vice president and general manager of the Data Center and AI Group, where she oversaw Xeon CPUs, GPUs, FPGAs and AI accelerators, and later led Altera through its spinout. In January 2026 she took the chair at VSORA, a French deep-tech company founded in 2015 that is moving its Jotunn8 inference processor into commercial roll-out.

Her argument, laid out in interviews this year, is specific. Customers are "increasingly looking for solutions to address the AI inference problem," she said. The economics she points to are unforgiving:

This is about power efficiency, cost per token, and certainly much lower latency and more determinism than you're able to get from more general purpose computing architectures.

The backdrop is a capex surge that has outrun even optimistic forecasts. Spending by Microsoft, Meta, Amazon and Alphabet reached $301 billion in the first half of 2026 alone, with the four guiding to $732.5 billion for the full year. Consensus for 2027 puts the quartet at $934.5 billion — Google $284.8 billion, Amazon $256.5 billion, Microsoft $207.6 billion, Meta $185.6 billion — within striking distance of a $1 trillion annual run rate.

That money is flowing into a semiconductor market that Gartner, in its August 2026 forecast, expects to grow 92% in 2026 to $1.56 trillion, with memory revenue more than tripling to $837 billion and AI data centers rising from 36.5% of industry revenue this year to more than 53% by 2030. IDC's base case projects $1.29 trillion in 2026 semiconductor revenue, up 52.8% from 2025, and a path to $1.75 trillion by 2030.

The question Rivera's outlook raises is whether this is a cyclical spending spike that will revert — or a structural shift in what kind of silicon the AI economy actually needs. The answer matters because it determines who wins: the general-purpose incumbents that dominated the training era, or the specialized inference vendors now coming to market.

Why Inference Economics Differ From Training

For the first phase of the AI boom, spending was dominated by training: giant clusters of GPUs running for weeks to produce foundation models. That phase is not over, but the center of gravity is moving. Once a model is trained, it is deployed — and deployed workloads run inference continuously, 24 hours a day, for every user query.

The economics differ sharply. Training is a one-off, latency-tolerant, throughput-driven workload where you can afford to burn power to finish a job. Inference is perpetual: every token generated carries a cost, every millisecond of latency affects user experience, and the bill arrives every month. Power efficiency and cost per token become the binding constraints.

Rivera frames VSORA's architecture around exactly this constraint:

We are focused on efficient data movement between the processing engine and external memory. We have a unique architecture that addresses the memory wall problem where compute units stall waiting for data, resulting in wasted performance and excessive power draw.

The "memory wall" is the mechanism, and it is worth understanding because it is the fulcrum of the whole thesis. Modern AI accelerators are so fast that they spend much of their time idle, waiting for weights and activations to arrive from memory. General-purpose GPUs were designed for graphics and parallel compute, not for the specific access patterns of large-language-model inference, where the same large weight matrices are read repeatedly and latency per token is the metric that matters. An architecture built from the ground up for inference — with large amounts of memory embedded close to the compute, deterministic latency, and specialized data paths — can deliver the same output with less energy and lower cost per query.

VSORA's Jotunn8 illustrates the claim: 3,200 teraflops at more than 50% utilization with 50% less power consumption than conventional GPUs, and 288GB of high-bandwidth memory at 8 TB/s throughput. The chip taped out in October 2025, entered manufacturing this year, and the company said in July that it is launching the commercial roll-out of the processor, backed by a funding round that brought Ardian into its shareholder base. The company works with TSMC and Global Unichip Corp to bring the processors from architecture to silicon.

The significance is not one startup's spec sheet. It is evidence of a broader industry movement toward heterogeneous architectures — purpose-built silicon for specific AI workloads rather than one general-purpose accelerator for everything. Rivera noted that "the investment by Nvidia in Groq certainly reinforces the position we've taken," validating the thesis that customers want inference-specific solutions. Nvidia agreed to a $20 billion chip licensing deal with Groq in December 2025 and unveiled an LPX platform incorporating Groq's inference technology; Groq itself raised $650 million in June 2026 to expand its inference cloud capacity. The message from the market leader is that inference is distinct enough to warrant its own architecture.

Cyclical Spending on Top of a Structural Shift

The cyclical-versus-structural question cuts two ways, and separating the two legs matters — because blending them produces the wrong verdict.

On the spending side, the capex cycle has cyclical fingerprints. Hyperscaler capital expenditure has historically moved in waves tied to product cycles, capacity utilization, and financing conditions. The 2026 surge — 158% above expectations from September 2024 — carries the hallmarks of a capacity race: companies building ahead of demand to secure supply of chips, power, and data-center space. If AI-generated revenue fails to keep pace with the infrastructure being laid down, the growth rate will decelerate. That is a cyclical risk, and it is real.

But the architecture shift underneath the spending is structural. Even if the growth rate of total capex slows, the mix of that spending is changing in a way that will not revert. Inference workloads will continue to grow as AI moves from experimentation into production across enterprises, consumer applications, automotive, and edge devices. Each deployed model creates a permanent, recurring inference load. The demand driver is not a one-time buildout but the cumulative stock of deployed AI — and that stock only grows.

Gartner's Ben Lee put it plainly: "The semiconductor industry is entering a fundamentally new phase of growth." The AI data center ecosystem's share of semiconductor revenue rising from 36.5% to more than 53% by 2030 is not a cycle; it is a reweighting of where value sits in the industry. Shrish Pant, also at Gartner, noted that "AI infrastructure has fundamentally changed the dynamics of the memory market," with DRAM revenue forecast to rise 246.6% in 2026 and NAND flash 371.9%.

The evidence floor for the structural call rests on three legs: a permanent change in workload composition (training giving way to persistent inference), a technological constraint that general-purpose architectures do not solve (the memory wall and power per token), and a reweighting of industry revenue toward AI data centers that forecasters expect to persist through 2030. The evidence floor for the cyclical call rests on capex growth rates that have more than doubled in a year and will almost certainly normalize.

The right read is both: a cyclical wave of spending riding on top of a structural shift in silicon demand. The cycle will determine the tempo; the structure will determine the winners.

Who Captures the Value — and the Counter-Case

The first-order reading of the capex numbers is simple: more spending is good for chip suppliers. The second-order question is which suppliers, and when the market reprices that distinction.

If inference efficiency becomes the binding constraint, the value pool shifts away from whoever sells the most raw teraflops toward whoever delivers the lowest cost per token. That favors specialized inference accelerators and custom ASICs over general-purpose GPUs for steady-state workloads. Data center discrete GPU revenue is expected to grow from $13.1 billion in 2023 to more than $51 billion in 2028, but revenue from custom AI ASICs deployed for inference is projected to grow faster from a smaller base.

The transmission chain runs: capex surge to buildout of inference capacity to power and cost-per-token constraints binding to heterogeneous architectures gaining share to memory and interconnect vendors benefiting regardless of which accelerator wins. Notice the last link: memory suppliers are a toll road on every architecture. Gartner expects memory to account for 54% of total semiconductor revenue in 2026, up from 27% in 2025. Whether the winner is a GPU, an inference ASIC, or a custom chip, they all need high-bandwidth memory and high-speed interconnect.

That is the cross-asset, cross-industry implication Rivera's outlook points to: the AI capex debate is often framed as a bet on one or two accelerator vendors, but the more durable exposure may sit in the infrastructure layer beneath them — memory, advanced packaging, power management, and optical interconnects.

There is also a geographic dimension worth weighing. VSORA's rise — and Ardian's investment — reflects a European push for "semiconductor sovereignty" in AI inference, a theme that runs alongside the U.S.-China technology competition. A French company taking on the inference market with backing from European investors and partnerships with Asian foundries is itself evidence that the AI chip landscape is fragmenting along both technical and geopolitical lines. That fragmentation creates opportunity for specialized players but also adds supply-chain and adoption risk.

The strongest case against Rivera's view is the simplest: this is a capex bubble, and when it deflates, specialized inference vendors will be the casualties, not the survivors.

The bear argument has teeth. Hyperscaler AI revenue, while growing fast, remains a fraction of the infrastructure being deployed on its behalf. If the return on AI investment disappoints, the first cuts will fall on unproven suppliers without established ecosystems. Nvidia's CUDA software moat, its scale, and its full-stack advantage mean that general-purpose GPUs remain the default choice; a startup with a better spec sheet but a smaller software stack faces a steep adoption curve. Gartner's Rajeev Rajput warned that "memflation will destroy, or at least delay, non-AI demand into 2028" — a reminder that the memory boom itself is a double-edged sword, raising costs across the industry and potentially choking off the very non-AI demand that would diversify the cycle.

There is also a timing risk. VSORA's Jotunn8 taped out in October 2025 and is entering commercial roll-out in 2026. If the capex cycle peaks before these chips reach scale, the company could arrive at a market that is suddenly more cautious about unproven silicon.

The falsifying signal is concrete: if hyperscaler capex growth decelerates by more than half year over year in 2027 — taking the four-company total well below the roughly $935 billion consensus — and if inference-specific accelerator revenue fails to grow faster than general-purpose GPU revenue over the same period, the structural-shift thesis is wrong and this was a cyclical spike after all. Watch the quarterly capex guidance from Microsoft, Alphabet, Amazon and Meta, and the revenue mix disclosures from the major accelerator vendors.

What to Watch: Scenarios Across Time Horizons

The base case is that AI infrastructure spending continues to grow through 2027, approaching $1 trillion annually across the largest hyperscalers, but at a slowing rate. Within that total, the share going to inference-optimized silicon rises. Beneficiaries split into two groups: the specialized inference vendors that can prove lower cost per token in real deployments, and the infrastructure layer — memory, advanced packaging, power and interconnect — that every architecture must buy.

The upside case: AI applications reach an inflection in enterprise adoption faster than expected, inference demand compounds, and purpose-built architectures capture meaningful share from general-purpose GPUs sooner than the market prices in. In that world, the current skepticism toward non-Nvidia AI chips looks like a buying opportunity.

The downside case: AI revenue fails to justify the infrastructure buildout, capex guidance is cut in 2027, and the market reprices the entire complex. Specialized vendors without scale or software ecosystems are the most exposed; memory suppliers face a sharper cyclical downturn as "memflation" reverses.

What to watch, by horizon:

  • Short term (next two quarters): hyperscaler capex guidance for the second half of 2026 and full-year 2027; any signs of guidance trimming.
  • Medium term (2027): benchmark results from inference accelerators, and customer deployment announcements that prove cost-per-token advantages.
  • Long term (through 2030): whether the AI data center share of semiconductor revenue moves toward Gartner's 53% forecast — the clearest read on whether this is a structural reweighting or a cyclical peak.

The central judgment: the AI capex number is the headline, but the architecture mix underneath it is the story. Rivera's bet is that the market is still pricing the buildout as if training-era economics apply — and that inference is where the next wave of value will be made.

The $700 billion question is not whether AI spending is too high — it is whether the market is buying the right chips for the job it is about to have.

Explore more exclusive insights at nextfin.ai.

Insights

What defines AI inference economics?

How does the memory wall limit GPUs?

Why is cost per token critical now?

Who leads 2026 AI capex spending?

Is AI spending cyclical or structural?

What is VSORA Jotunn8 chip architecture?

How does Nvidia view inference markets?

Why does memory revenue triple 2026?

What risks face inference chip vendors?

Why Europe seeks chip sovereignty?

What signals a capex bubble bursting?

Who benefits AI memory demand surge?

Can startups beat Nvidia CUDA moat?

How do AI training and inference differ?

Where does AI silicon value shift?

2027 hyperscaler capex forecast details?

Why power efficiency key for inference?

Why Groq deal matters for inference?

What if capex growth slows sharply?

Who wins specialized vs general chips?

Search
NextFinNextFin
NextFin.Al
No Noise, only Signal.
Open App