NextFin News - OpenAI on Tuesday released 722 manuscripts covering at least 372 long-standing mathematical problems - and the most unsettling part of the announcement is not what the AI proved, but what it implies about the business of selling AI. A month after the same lab used a roughly 10,000-agent swarm costing millions of dollars in computing power to crack the Navier-Stokes Millennium Prize Problem, nearly every one of the 372 follow-up results came from a single prompt handed to a single agent, at an average cost of about three hours of ChatGPT Pro thinking time. The gap between those two numbers is where the token-selling business model starts to break.
The Deluge: 372 Result Families, One Prompt Each
The sequence matters. On September 8, mathematicians at OpenAI announced that an internal system, significantly more capable than the publicly known GPT-6 Astra, had found a finite-time singularity in the three-dimensional Navier-Stokes equations - resolving one of the six remaining Millennium Prize Problems posed in 2000 by the Clay Mathematics Institute, each carrying a $1 million prize. That proof, produced with the collective effort of roughly 10,000 autonomous agents running for approximately 50 hours, was formally checked in the Lean proof language and stood as the most expensive AI-generated theorem in history.
On October 6, at 6 P.M. Eastern, the company published the new results in a public GitHub repository under an Apache 2.0 license: 722 manuscripts organized into 372 result families across 17 fields, including algebra, number theory, theoretical computer science, mathematical logic, and topology. The catalog includes a claimed solution to the four-dimensional Kakeya conjecture, a new zero-free region for the Riemann zeta function for Re(s) greater than 11/12, progress on the Hodge Conjecture for CM abelian varieties, and resolutions of long-standing Erdős problems in extremal graph theory and multicolor Ramsey numbers. The company said it posed approximately 4,000 problems to the model in total, aggregating the output into result families and requiring an appropriate level of significance before publication. According to the repository's formalization catalogue, 162 of the 722 papers carry a Lean-formalized main result.
The cost collapse is the headline within the headline. OpenAI published 10 abridged summaries of the model's reasoning, compute estimates expressed in ChatGPT Pro usage, and statistics on attempted problems. The average result used the equivalent of roughly three hours of ChatGPT Pro thinking. In other words, a result set that would have consumed a small army of researchers for years was produced at a consumer-subscription price point - and almost none of it required the million-dollar swarm architecture that produced Navier-Stokes.
"We're releasing a broad range of new mathematical results produced by an internal frontier model," the company said in its announcement, adding that it is "working to responsibly release the model that produced these results."
The release was shaped in consultation with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, which had published responsible-release recommendations on September 29 after receiving more than 600 replies from the mathematical community. The company said it drew on that advice for this release - with one conspicuous exception. The advisory group had asked AI labs to stop testing frontier mathematical capabilities on proprietary, unreleased models. OpenAI did the opposite: it published the outputs while withholding the model itself.
The Scary Implication: Why the Labs May Exit the Token Business
Here is the implication that should worry anyone holding AI infrastructure stocks on the assumption that token sales are the endgame. For three years, the dominant business model in frontier AI has been straightforward: train a massive model, then rent it out by the token through an API. Every dollar of revenue is a function of how much inference the market consumes. That model depends on the model being a sellable commodity - valuable enough that customers pay for access, but not so valuable that handing it out undermines the owner's advantage.
The math results puncture that balance. When a single prompt to a single agent can produce hundreds of frontier research results, the model is no longer a product you want to rent. It is a proprietary research institution - a capability so concentrated that selling access to it by the token means selling away the advantage itself. The rational move is to stop selling the model and start selling what the model builds: products, enterprise solutions, and credibility signals that competitors cannot replicate because they cannot inspect the engine.
This is not speculation about intent; it is an inference from behavior. The company has not released the model. It has released the outputs, the reasoning traces, and the compute estimates - everything a researcher needs to believe the capability is real, and nothing a competitor needs to reproduce it. The release functions as a capability signal aimed at researchers, rival labs, and enterprise customers deciding whose reasoning systems to build on. Math benchmarks have become one of the clearest, hardest-to-fake proxies for raw reasoning capability in frontier AI. Publishing the proofs without the model is how you win that proxy war without arming your rivals.
The financial pressure points in the same direction. The price of producing GPT-4-quality output has fallen from about $30 per million input tokens at the model's March 2023 launch to under $0.50 per million tokens for equivalent capability today - a decline of more than 98% in roughly three years, driven by algorithmic efficiency, cheaper hardware, and open-weight competition. Meanwhile, internal financial documents reported in the press point to a projected operating loss of roughly $14 billion in 2026 and cumulative losses of about $44 billion through the end of the decade, even as the company's revenue run rate approaches $20 billion a year. When the price of your product collapses toward commodity levels and your capital bill keeps rising, the only escape is to move up the value chain - to sell the finished good, not the raw input.
Even the company's chief executive has acknowledged the tension. In an interview this week, he said he wants to put the technology "in everyone's hands," while adding that "concentration of power with AI is a terrifying thing." The two statements cannot both be fully honored. The math release suggests which one is winning: the capability stays concentrated, and the world gets to see what it can do.
The Second-Order Effect: Benchmarks Become Weapons, Not Public Goods
The first-order reading of this news is the one already priced into AI chip valuations: AI is getting dramatically smarter, so the compute buildout continues. That reading is not wrong, but it is incomplete, and it is the incompleteness that matters for investors.
The second-order effect runs through the revenue line that justifies the chip spending in the first place. The hyperscalers and chipmakers are betting on a world in which frontier inference is consumed in ever-larger volumes through open APIs. If the labs that own the frontier models conclude that those models are too valuable to sell, the token market stops growing at the rate the infrastructure buildout assumes. Value migrates from the commodity input - tokens, GPU hours - to the application layer that owns the customer relationship and the proprietary model underneath it.
There is a third-order consequence, and it lands on the mathematics community itself. The company has acknowledged that many of the newly published results are not yet understood even by its own mathematicians. A field whose currency is comprehension is being asked to validate outputs faster than humans can read them. Formal verification in Lean can certify that a proof is logically correct; it cannot certify that the proof contains a novel idea worth building on, or that it did not assemble existing techniques in a way that sidesteps academic norms. The advisory group's request that labs stop frontier testing on closed models was precisely an attempt to preserve the field's role as validator. Its rejection signals that the pace of production has already outpaced the pace of governance.
Is this a cyclical fluctuation or a structural shift? It is structural, and the evidence is in the cost curve rather than the headlines. A cyclical claim would require a short-term driver - a supply shock, a liquidity squeeze, an inventory build - that mean-reverts on its own. None of those applies here. What changed is the production function for mathematical knowledge itself: the marginal cost of generating a frontier-level result has fallen from years of specialist labor to hours of consumer-subscription compute, and that cost does not bounce back. The regime change is technological and institutional - a new class of agentic systems that will not self-correct into being less capable - and it is reinforced by the financial pressure on the labs to monetize capability rather than tokens. The short-term leg is cyclical in one narrow sense: the credibility signal from this specific release will fade as the community vets the manuscripts, and any retraction would temporarily restore the field's bargaining power. But the long-term leg is structural: even a full retraction of every result would not restore the pre-agentic cost of discovery. The capability, once demonstrated, cannot be un-demonstrated, and the next release will start from a higher baseline.
This is the real asymmetry of the moment. The lab that controls the model controls both the rate of discovery and the terms of disclosure. Everyone else - rival labs, academic mathematicians, investors in the infrastructure stack - is reacting to a signal they cannot independently reproduce.
The Counter-Thesis: Why the Token Business Is Not Dead Yet
The strongest case against this reading is simple: the labs still need the revenue. Renting models by the token is the only business line that currently generates tens of billions of dollars a year, and that cash funds the training runs that keep the lead intact. Abandoning it would be financial suicide. Open-source competition from Chinese labs and profit-insensitive initiatives from the large platforms keeps pressure on any lab that tries to wall off its models - if you do not sell, someone else's open release will commoditize you anyway. And the chief executive's stated preference for broad distribution is not obviously a lie; concentration of power is a genuine safety and reputational risk.
This counter-thesis has force, but it conflates the transition with the destination. The token business does not need to die tomorrow for the thesis to hold; it needs only to stop being the growth engine. A lab can keep selling API access to yesterday's models - the ones already commoditized - while reserving the frontier for product and enterprise use. That is not an exit; it is a stratification, and it produces the same terminal outcome: the highest-margin capability is no longer for sale by the token.
The falsifying signal is concrete. If the company releases the weights or full API access to the model that produced these 372 results within the next six months, or if API revenue from frontier models continues to grow at a rate consistent with the infrastructure buildout assumptions, the "exit the token business" thesis is wrong. Watch the release status of the model and the revenue mix in the next earnings disclosures from the major cloud platforms. A shift toward enterprise product revenue and away from raw token consumption is the confirmation.
Who Benefits, Who Is Exposed
The beneficiaries of this shift are not the obvious ones. In the short term, the narrative still favors the chipmakers and the hyperscalers - the math results are read as proof that the AI arms race is accelerating, and private-market valuations reflect it, with the lab's own valuation reported near $500 billion after a secondary share sale. But the structural reading favors a different set of winners: the application-layer companies that can embed proprietary reasoning into products with defensible margins, and the labs that can convert capability signals into enterprise contracts rather than commodity API volume.
The exposed are the investors who have underwritten the infrastructure buildout on the assumption that token consumption will scale linearly with capability. If the frontier models become too valuable to rent, the consumption curve flattens even as capability rises - a divergence that the current multiples do not price. It also exposes the academic mathematics community, whose role as the field's validator is being bypassed by a combination of computer-verified proofs and corporate-controlled disclosure.
What to Watch: Three Horizons
Short term (weeks): The mathematics community parses the 722 manuscripts. Expect a wave of scrutiny over novelty and attribution - the same frustration that followed the Navier-Stokes result, when some experts concluded the AI had assembled existing ideas rather than generated new ones. Any high-profile retraction or misconduct finding would dent the credibility signal and could push the lab toward more openness.
Medium term (months): The model release decision. The company has said it is "working to responsibly release the model." If that release is delayed, narrowed, or converted into an enterprise-only product, the token-exit thesis strengthens. The advisory group's next public statement - and whether other labs follow its guidelines - is the governance signal to watch.
Long term (years): The business-model settlement. Either frontier AI becomes a product layer, with labs monetizing applications and keeping models closed, or it remains a commodity input sold by the token. The two paths imply very different winners: the first concentrates value at the lab and application layer; the second spreads it across the infrastructure stack.
The base case is stratification: yesterday's models sold as cheap tokens, frontier models reserved for products. The upside case for the infrastructure bulls is that capability growth is so explosive that even closed frontier models generate enough product revenue to justify ever-larger compute budgets. The downside case is a demand gap - capability races ahead of monetization, and the capex cycle overshoots the revenue that can actually be captured.
One month ago, the story was that AI had solved one of the hardest problems in mathematics. This week, the story is that it solved 372 more before breakfast - and in doing so, may have solved the AI labs out of the business they thought they were in. The scary implication is not that the machines are too smart to sell. It is that they are too smart to sell cheaply, and everyone who priced the future on cheap tokens is now holding a different asset than they thought.
Explore more exclusive insights at nextfin.ai.
