NextFin

Chinese AI Models Enter a Performance-to-Cost Race

Summarized by NextFin AI
  • Chinese AI competition is shifting from raw model capability toward performance-to-cost efficiency as inference becomes cheaper and open-weight models spread internationally.
  • Chinese inference costs have fallen by more than 95%, lowering experimentation barriers and potentially expanding demand for cloud infrastructure, enterprise software, and AI applications.
  • Open-weight distribution accelerates adoption but weakens direct monetization, forcing developers to explore commercial licensing, hosted APIs, subscriptions, and cloud services.
  • The investment outcome remains unsettled: cheap models may drive broad ecosystem growth, but sustained paid usage and recurring revenue must prove that adoption converts into durable supplier economics.

NextFin News - The most important claim in Goldman Sachs’ August 5 view on Chinese artificial intelligence is not that another model has reached a new benchmark. It is that the competitive prize is moving from raw capability to the best performance-to-cost balance. Ronald Keung, the bank’s head of Asia internet research, said Chinese developers should compete more intensely on that measure as inference becomes cheaper and open-weight models spread beyond their home market.

That distinction matters for investors because a cheaper model can produce two opposite outcomes. It can expand demand for cloud computing, applications and enterprise software, or it can commoditize the model layer and leave developers with little pricing power. The evidence so far points to a structural change in how AI is distributed, but not yet to a settled business model.

Goldman Sachs said Chinese inference costs had fallen by more than 95% over the prior year. In the same published discussion, Keung cited a Chinese model priced at $0.14 per million input tokens, a level he described as only a single-digit percentage of the price charged by an equivalent reasoning model from a large US technology company. The implication is straightforward: if the cost of asking a model to perform a task collapses, more companies can afford to use it, and developers can put AI inside more products.

The second implication is harder. Open-weight distribution allows third parties to download, modify and host a model’s parameters. That increases reach but can sever the link between usage and the original developer’s revenue. Keung has therefore pointed to commercial licensing for providers that host open-weight models as one possible way for Chinese developers to capture value. The market is not deciding whether Chinese models are useful. It is deciding who captures the economics once usefulness becomes cheap.

The August 5 video discussion places that question in a broader competitive frame. Chinese labs are no longer competing only to close a capability gap with US frontier models. They are competing for developers, enterprise workloads and the feedback generated by real-world use. The first phase of the race rewarded model launches. The next phase will test adoption, retention, compute efficiency and monetization.

The Headline Is Cost, but the Mechanism Is Adoption

The direct story is that lower inference prices make Chinese models more competitive. The deeper mechanism is a fall in the fixed cost of experimentation, which changes who can try AI and how often they can use it.

Inference is the operating stage after training, when a model responds to new inputs. Training attracts headlines because it requires large pools of chips and capital. Inference determines whether a model can be embedded in a customer-service workflow, a coding tool, a search product or an industrial process at a price a customer will accept. A 95% reduction in that operating cost does not automatically create revenue, but it changes the break-even point for thousands of applications.

Suppose an enterprise previously treated a model call as an expensive exception. A lower token price permits more testing, longer context, more frequent automation and a wider range of employees or customers using the system. That produces a demand response before it produces a margin response. The model developer may earn less per token while the ecosystem earns more in aggregate.

Keung’s $0.14 per million input-token example captures the asymmetry. At that price, a developer can run many more experiments before the model bill becomes material. DeepSeek’s official API documentation, as of the data cutoff, listed $0.14 per million input tokens for cache misses on its V4 Flash model and $0.28 per million output tokens. Alibaba Cloud’s official pricing documentation listed $0.10 per million input tokens and $0.40 per million output tokens for Qwen3.5 Flash for international deployment. Goldman’s $0.14 example and these current vendor schedules are not necessarily the same model or commercial configuration. They nevertheless demonstrate the same direction: low-cost Chinese offerings are making usage economics a central competitive weapon.

The important cross-industry transmission is from model prices to cloud utilization. A model that is inexpensive to call can drive more API traffic, but the traffic still needs storage, networking, orchestration, security and often specialized compute. Cloud providers can therefore gain even if the model itself is priced aggressively. That is why the model layer and the infrastructure layer should not be treated as the same business.

The second-order effect reaches application companies. Cheaper inference lowers the cost of adding AI features, but it also lowers the barrier for competitors to copy those features. The winner may be the company with proprietary distribution, workflow data or customer relationships rather than the company with the lowest model price. AI can raise software demand while compressing the value of undifferentiated software features.

That is the first expectation gap. Investors who focus only on model quality may miss that price is becoming the product feature. Investors who focus only on price may miss that the revenue pool could migrate into cloud and applications.

Why Open Weights Expand Reach and Complicate Revenue

Open-weight distribution is the structural feature that makes the cost story more powerful and the monetization story more difficult.

When a developer releases model weights under a permissive license, users can download and run the model on their own infrastructure or through another cloud provider. DeepSeek’s public model documentation says its V3 model supports commercial use, and the published model license grants broad, no-charge rights to reproduce, distribute and create derivatives of the model, subject to the license terms. That model can travel through a developer ecosystem without every use passing through the original company’s API.

The benefit is speed. An enterprise with data-residency requirements can self-host. A startup can fine-tune the model for a narrow task. A cloud platform can expose the model alongside competing models. Every deployment can improve familiarity and generate application-level demand even when the originator does not collect a conventional per-token fee.

The cost is revenue leakage. A model developer can spend heavily on research and training while a third-party host captures the usage payment. The more capable and permissive the model, the more credible the risk that model intelligence becomes an input purchased once and reused many times.

Keung’s commercial-licensing idea addresses that gap. A developer could leave weights available for research or limited use while requiring commercial providers to buy a license to serve the model at scale on their own infrastructure. That would convert distribution into a negotiated business relationship. It would also introduce friction that could reduce adoption, especially where a rival model remains free and sufficiently good.

“What’s clear to us is that lowering the cost of AI models will drive much higher adoption, as it would make the models much cheaper to use in future,” Ronald Keung, head of Asia internet research at Goldman Sachs, said in the bank’s published discussion.

The commercial question is not whether licensing is possible. It is whether the developer can impose it without giving up the network effects created by openness. A restrictive license may capture more revenue per customer but shrink the number of developers willing to build around the model. A permissive license may maximize adoption but shift the economics toward cloud hosting, implementation services and applications.

This creates a familiar technology tradeoff in a new form. Distribution builds the ecosystem; scarcity supports margins. Chinese developers are testing whether they can have both through a tiered model: open access for visibility, paid weights or hosted APIs for commercial scale, and subscription products for end users.

The second-order transmission runs through bargaining power. If several Chinese models offer comparable performance and similar prices, cloud platforms can switch providers and customers can multi-home. That weakens the model developer’s leverage. If one model develops a clear lead in coding, reasoning, Chinese-language enterprise data or agentic workflows, the same open ecosystem can become a funnel into higher-value paid services.

The measure to watch is therefore not downloads alone. It is the conversion from open usage to recurring commercial revenue, including subscription revenue, API traffic, licensing fees and cloud consumption. Adoption without conversion can still matter strategically, but it does not validate a high-margin model-company thesis.

Cyclical Excitement Meets a Structural Cost Reset

The current enthusiasm around Chinese AI has a cyclical component, but the cost reset is structural.

The cyclical component is the model-release cycle. Each new launch can attract attention, lift expectations for related technology names and prompt a burst of benchmark comparisons. That attention can fade when users discover that benchmark performance does not translate into reliable production workloads, when capacity constraints appear, or when a new release makes the previous one economically obsolete. Model rankings and technology valuations can mean-revert.

Three historical comparisons explain why caution is necessary. The first is the smartphone market, where rapid hardware improvement eventually shifted value away from device specifications and toward operating systems, distribution and services. The second is cloud computing, where lower unit costs expanded usage but forced providers to compete on scale, uptime and integrated services rather than raw compute alone. The third is internet search and advertising, where more usage did not guarantee that every technology supplier captured the resulting economic value.

These comparisons share a mean-reversion pattern: early scarcity produces high margins and high valuation narratives; capacity and competition then reduce unit prices; durable winners emerge only where they control distribution, data, infrastructure or customer workflows. The AI model cycle can follow the same path.

But the structural component is different. Inference costs falling by more than 95% in one year, open-weight releases that permit commercial deployment, and the spread of model competition across Chinese developers alter the market’s operating rules. Those forces do not reverse merely because the next model disappoints. Even if model quality improves more slowly, enterprises that have built workflows around low-cost inference will have an incentive to keep those workflows and optimize them further.

That makes the proper call a split verdict: model excitement is cyclical; low-cost, widely distributed inference is structural. The mistake would be to treat a temporary valuation wave as proof of permanent pricing power, or to treat falling model prices as proof that the entire AI economy is unprofitable.

Goldman’s official research also describes Chinese hyperscalers as beginning to rely on home-grown AI chips and hardware while serving mainly domestic customers, with subscription revenue already emerging from AI applications. That combination could make lower model costs more useful within China, although it does not remove hardware constraints or prove that deployment will be profitable.

That resilience has a limit. Hardware availability, energy, data-center construction, regulatory approval and the quality of enterprise integration still constrain actual deployment. A low token price is an invitation to use AI, not proof that a customer can deploy it securely, quickly and profitably.

The Counter-Thesis: Cheap Models Could Destroy the Investment Case

The strongest counter-thesis is that Chinese AI developers are entering a commodity market before they have built durable monetization. If model capability converges and inference prices continue to fall, customers may capture most of the benefit while developers absorb the cost. In that case, the 95% cost decline is not an adoption flywheel for model companies; it is a margin collapse disguised as progress.

This argument has three foundations. First, open weights make substitution easy. A customer can test multiple models, fine-tune one internally and change providers without rebuilding an entire application. Second, the largest cloud platforms have incentives to subsidize model access to sell infrastructure and data services. Third, benchmark leadership can be temporary because rivals can reproduce techniques, distill models or release a cheaper version.

The model-layer bear case is therefore not a straw man. It is consistent with the history of computing markets, where falling unit costs often create enormous social value while narrowing supplier margins. It also explains why licensing could prove harder than expected: a commercial restriction that protects revenue may be less attractive to users than a permissive rival.

The bullish answer is that commoditization at the model layer can be constructive for the wider AI stack. Lower prices increase experimentation, and experimentation creates demand for compute, networking, security, workflow software and specialized applications. The original model developer can still participate through cloud distribution, premium versions, enterprise support, subscriptions or licensing. The business becomes less like selling a scarce digital product and more like operating an ecosystem.

That answer is plausible but conditional. It requires evidence that usage is converting into recurring revenue and that cloud traffic is incremental rather than merely shifting from one provider to another. It also requires developers to retain enough technical differentiation to prevent every new model from becoming interchangeable.

The falsifying signal for the structural-adoption thesis would be concrete: if Chinese model API traffic and enterprise subscriptions fail to grow for two consecutive quarters after major price cuts, while paid licensing remains negligible, lower cost will have behaved mainly as a transfer of value to users. The falsifying signal for the model-layer bear case would be equally concrete: if a leading Chinese developer reports sustained growth in paid API, subscription or licensing revenue for two consecutive quarters while maintaining open-weight distribution, the ecosystem model will have demonstrated conversion rather than just reach.

Keung’s public comments support the adoption case, but they do not settle the conversion case. He has said domestic growth is very fast and that small and medium-sized enterprises globally, along with larger companies, are beginning to consider Chinese models. Consideration is not deployment, and deployment is not revenue. The next proof must come from operating metrics.

What the Race Means for Markets and Companies

In the short term, the performance-to-cost race is likely to favor companies that can demonstrate model releases, user growth or access to cloud distribution. That is the sentiment and liquidity horizon. It can lift expectations quickly, but those expectations are vulnerable to benchmark reversals, capacity problems and evidence that usage is not monetizing.

In the medium term, the beneficiaries should be more specific. Cloud platforms can gain from higher inference volumes and from enterprise demand for hosting, security and orchestration. Application companies with proprietary distribution can use lower model costs to improve margins or add functionality. Chip and data-center suppliers benefit only if rising utilization offsets lower revenue per model call and if local supply constraints do not prevent deployment.

The exposed companies are those whose only advantage is access to a model that rivals can copy or download. They face a double pressure: customers demand lower prices while competitors gain access to similar capabilities. A model company that cannot turn openness into cloud, licensing or subscription revenue may create a valuable ecosystem for someone else.

The long-term structural implication is a more plural AI market. Open-weight Chinese models can spread through global developers even when commercial access, data governance or geopolitics limits direct cross-border deployment. That does not mean Chinese models will replace US frontier systems. It means the frontier may fragment by task, price, language, sovereignty and infrastructure. A single global ranking will become less useful than the economics of each workload.

The base case is a two-speed market. Low-end and routine workloads move rapidly toward cheap models, while the hardest reasoning, regulated enterprise applications and highly integrated agents retain a premium. In this scenario, cloud and application revenue grows faster than model pricing, and developers capture value through a bundle of API, subscription and licensing channels. The trigger would be two consecutive quarters of rising paid usage alongside stable or improving customer retention.

The upside case is a faster adoption loop. Chinese models become good enough across coding, reasoning and multimodal tasks that lower cost unlocks a large population of small and medium-sized enterprises. Cloud utilization, application subscriptions and model licensing rise together. The trigger would be broad evidence of international enterprise deployments and recurring revenue, not another benchmark win.

The downside case is a price war without conversion. Developers cut prices, release weights and subsidize APIs, but customers do not expand production workloads because of security, reliability, data-residency or integration barriers. Model margins fall, cloud demand remains incremental rather than new, and the valuation premium attached to AI names contracts. The trigger would be two quarters of flat enterprise usage despite further price reductions.

For investors, the key distinction is between the technology’s social surplus and the supplier’s financial surplus. Goldman’s thesis is strongest on the first: cheaper models should broaden adoption. The unresolved question is where the second accumulates.

The August 5 discussion therefore marks a transition in the story. Chinese AI developers are moving from proving that they can build capable models to proving that they can build durable businesses around cheap ones. That is a harder test because the same efficiency that attracts users can remove the scarcity that supports pricing power.

Cheap inference is not the end of the AI investment case. It is the end of the simple one.

Explore more exclusive insights at nextfin.ai.

Insights

What does performance-to-cost mean in the Chinese AI model market?

How does inference differ from AI model training?

Why have Chinese AI inference costs fallen by more than 95 percent?

How do lower inference prices encourage enterprise AI adoption?

What advantages do open-weight Chinese AI models offer developers?

Why can open-weight distribution weaken model developers’ pricing power?

How could commercial licensing help Chinese AI developers monetize open models?

What role will cloud providers play in the Chinese AI cost race?

Which companies could benefit most from cheaper AI inference?

Why might lower model prices increase application revenue while reducing model margins?

How do Chinese AI models compare with equivalent US reasoning models on price?

What lessons do smartphones, cloud computing, and search offer for AI markets?

What challenges could prevent cheap AI models from reaching production workloads?

How might data residency, regulation, and geopolitics shape Chinese model adoption?

Could Chinese AI model competition develop into a commodity price war?

What operating metrics would prove that low-cost AI usage is becoming recurring revenue?

How could the AI market fragment by task, language, price, and infrastructure?

What are the likely bullish and bearish scenarios for Chinese AI developers?

Search
NextFinNextFin
NextFin.Al
No Noise, only Signal.
Open App