NextFin News - Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash on September 10, 2026, a cheaper, faster model that it says outperforms its own higher-priced V4 Pro on key agentic and coding benchmarks - and the market immediately asked the question Wall Street has been dodging all year: if frontier capability keeps getting cheaper, what happens to the trillions of dollars of compute spending propping up the AI trade. The answer showed up fast. Shares of Chinese AI rivals MiniMax Group and Z.AI plunged more than 8% in Hong Kong, Alibaba slid more than 2%, and memory makers Samsung Electronics and SK Hynix each dipped more than 3% the next day, while Nvidia closed September 10 at $218.36, down 2.37%, as chip stocks endured an ugly session.
The release is not a one-off price cut. It is the latest move in a price war DeepSeek started and has been winning, and it lands as the Hangzhou-based startup prepares an initial public offering on Shanghai's STAR Market. The combination is what rattled investors: a company that is both compressing the economics of AI inference and racing to the public markets, with a valuation that remains a fraction of the sums being sought by OpenAI and Anthropic.
The Release: A Cheaper Model That Claims to Beat the Flagship
DeepSeek said the new model is designed for greater capability, faster inference, higher throughput and scaling to larger models, according to its statement. The specifications tell a more pointed story. V4.1 Flash carries 552 billion total parameters but activates only 8 billion on prefill and 16 billion on decode, using a new causal encoder-decoder architecture with 40 layers, 384 routed experts, and a native 1-million-token context window. It handles both text and image input, ships under an MIT license, and costs $0.15 per million input tokens and $0.60 per million output tokens during off-peak hours - with cache hits priced at just $0.003.
On DeepSeek's own benchmark table, V4.1 Flash beats V4 Pro on DeepSWE v1.1 (74.2 versus 62.7) and Terminal-Bench 2.1 (90.6 versus 87.9), while losing on the HLE knowledge benchmark (36.8 versus 42.7). The practical consequence is severe for anyone selling expensive inference: starting at noon Beijing time on September 14, DeepSeek will retire the V4 Pro endpoint and route every request to deepseek-v4-pro through V4.1 Flash, billed at Flash rates. That is a 4.4x cut on peak input pricing and a 3.3x cut on peak output pricing for existing V4 Pro users - a revenue hit wrapped in a capability claim.
This is not DeepSeek's first strike. The company made a 75% discount on V4 Pro permanent in May, dropping the price from $3.48 to $0.87 per million tokens, then raised some rates in August before this week's reset. The direction of travel is unambiguous: the floor keeps falling, and DeepSeek keeps setting it.
Why the Market Flinched: The Compute Thesis Is Being Priced for Scarcity That May Not Exist
The chip rally of 2025 and 2026 was built on a simple assumption: frontier AI requires ever more GPUs, ever more high-bandwidth memory, and ever more power, so the suppliers of those inputs capture the value. DeepSeek's model family attacks the first link in that chain. V4.1 Flash is a mixture-of-experts model that activates a small fraction of its parameters per token, and it is explicitly engineered for inference efficiency - the exact workload that consumes the bulk of deployed compute. When the marginal unit of AI output needs fewer tokens, fewer GPUs, and less memory, the demand curve that justified record capital spending softens even if the number of users grows.
The efficiency gains are not theoretical. Independent research firm Artificial Analysis measured the cost of completing its Intelligence Index test battery across major models and found V4-Flash ran at roughly three cents per test. The nearest comparisons were not close: Moonshot's Kimi K3 cost 86 cents, OpenAI's GPT-5.6 Sol cost $1.86, and Anthropic's Claude Fable 5 cost $3.15. On the same firm's composite, V4-Flash scored 50 out of 100, tying Google's Gemini 3.6 Flash and sitting a single point below Meta's Muse Spark 1.1 and Z.AI's GLM-5.2 - while Kimi K3 reached 57. In other words, DeepSeek is delivering mid-pack-to-strong performance at roughly one-fiftieth the price of the premium tier.
That math is what hit memory makers on Friday. Samsung Electronics and SK Hynix, which together account for roughly half the weight of South Korea's Kospi index and have been the primary beneficiaries of the high-bandwidth memory shortage, each fell more than 3%, paring their rebound from July's steep selloff. The logic is direct: if each unit of AI output requires fewer tokens, fewer GPUs, and less memory, the demand curve that justified record HBM investment softens. Reports that DeepSeek is developing its own AI chip added a second layer of pressure - the customer may become the competitor.
The broader chip complex was already fragile before the announcement. Broadcom plunged roughly 15% on September 10 after a weaker-than-expected AI chip outlook and a decision to reiterate rather than raise its 2026 guidance, and Nvidia's market capitalization stood at about $4.97 trillion - a level that prices in years of uninterrupted demand growth. DeepSeek did not need to move the needle much to remind investors how much of that valuation rests on assumptions about compute intensity.
The Transmission Chain: From a Price Cut to a Multiple Compression
The market's reaction looks outsized for a single startup's model launch until the chain is traced. The first-order effect is the most visible: DeepSeek's API pricing sets a new reference point, and customers with elastic workloads migrate. OpenRouter data shows the migration is already well advanced - Chinese-origin models grew from under 2% of token consumption in late 2024 to more than 50% by June 2026, and DeepSeek V4 Pro at $0.87 per million output tokens runs at one-fifty-seventh the price of Anthropic's Fable 5.
The second-order effect is where the damage compounds. When the reference price for a unit of capability falls, every buyer of AI - from a startup to a hyperscaler - recalculates what they are willing to pay. That is why Uber exhausted its 2026 AI budget in four months and why Microsoft reduced Claude Code access internally: spending is being rationed against a price benchmark that keeps sliding. The pressure does not stop at the model layer. Cheaper inference means fewer GPUs per query, less high-bandwidth memory per query, and less power per query - which is why the selloff leaked from AI labs to memory makers to the entire chip complex.
The third-order effect is the expectation gap. Hyperscaler valuations and capital expenditure plans were underwritten on a world in which compute demand grew faster than efficiency improved. DeepSeek's release forces investors to ask whether that assumption still holds. If efficiency gains outrun usage growth, the suppliers of compute face a volume-versus-price squeeze: more queries, but less revenue per query and less hardware per query. That is the scenario the September 10-11 selloff priced in, at the margin.
The IPO Subplot: A Valuation Gap That Cuts Both Ways
The timing matters as much as the technology. DeepSeek has engaged CITIC Securities to prepare for an initial public offering on Shanghai's STAR Market and aims to begin the process this year, according to people with knowledge of the matter. The company is raising a new funding round at a valuation of 500 billion yuan, or about $75 billion, after closing roughly $7.4 billion in June at a post-money valuation exceeding $50 billion, with Liang Wenfeng personally contributing 20 billion yuan and Tencent Holdings and CATL adding 10 billion yuan and 5 billion yuan respectively.
That valuation looks cheap next to U.S. rivals. Some investors expect Anthropic's IPO to carry a valuation as high as $2 trillion, and OpenAI is targeting as much as $1 trillion. The gap is not just geography - it reflects a stark difference in revenues, and how hard it has been for Chinese AI companies to turn technical progress into income. But it also frames the central tension: if DeepSeek can list at $75 billion while pricing inference at a fraction of the U.S. rate, the market is being asked to believe that the American labs' premium pricing is defensible.
There is a reason the pricing pressure is existential for the U.S. labs. OpenAI reported a negative adjusted operating margin of 122% in the first quarter of 2026, according to reporting on its financials, and Anthropic's first profitable quarter arrived only in the second quarter of 2026, driven almost entirely by Claude Code. Premium API pricing is what services their capital-raising. DeepSeek's price cuts do not just take market share - they shrink the margin story that justifies the valuations.
The Counter-Thesis: Demand Elasticity and the Sovereign Compute Moat
The strongest case against the bearish read is simple: cheap AI does not shrink compute demand, it explodes it. Goldman Sachs expects agentic AI token consumption to grow roughly 24 times by 2030. If the price of inference falls by 50x, usage does not stay flat - it expands, and the suppliers of GPUs, memory, and power still win on volume. Enterprise behavior supports this: Uber exhausted its 2026 AI budget in four months, and the shift to Chinese models is adoption accelerating into the price cuts, not away from them. Under this view, the September 10-11 move was a cyclical overreaction to a structural demand story that remains intact.
There is also a moat argument. Daniel Morgan, a senior portfolio manager at Synovus Trust Company, called the selloff an over-reaction, noting that DeepSeek's models compete with ChatGPT, Meta, and Alphabet on the application layer.
The real money in AI is providing the chips for the data centers.
The hyperscalers are not passive: Google has agreed to rent computing capacity from SpaceX, paying $920 million per month for 110,000 Nvidia GPUs and related components from October 2026 through June 2029. OpenAI is deepening cooperation with Samsung Electronics as it develops its own chips, expanding ties across semiconductors and enterprise AI services. The infrastructure layer is consolidating, not fragmenting.
And the capability gap has not closed. On knowledge-intensive benchmarks like HLE, V4.1 Flash still trails V4 Pro. Enterprise customers with compliance, latency, and integration requirements do not switch on price alone, and the U.S. labs retain advantages in distribution, trust, and the developer ecosystems built around their platforms. The January 27, 2025 DeepSeek shock - when Nvidia fell 16.9% and erased about $593 billion in a single session, the largest one-day loss for a Wall Street stock on record - proved the market can overreact to a Chinese efficiency breakthrough. That selloff reversed.
But this time is different in one crucial respect. In January 2025, DeepSeek-R1 was a surprise - a reasoning model trained at a fraction of the expected cost. By September 2026, the market has had months to absorb the implication. Lian Jye Su, chief analyst at Omdia, said of the April V4 launch that "this announcement followed a rather predictable path." The surprise has worn off; what remains is the grind. A one-day shock can reverse. A structural compression of pricing power cannot be undone by a rebound.
What to Watch: The Signal That Decides Which Thesis Is Right
The falsifying signal is concrete. If hyperscaler capital expenditure guidance for 2027 holds or rises through the next two earnings seasons - and if high-bandwidth memory contract pricing stays firm through the first half of 2027 - the demand-elasticity thesis wins, and the September selloff was a cyclical overreaction. If, instead, capex guidance is cut while inference volumes continue to grow, the efficiency thesis is confirmed: the market is paying for compute intensity that no longer exists, and the multiple compression has further to run.
For the AI labs, watch the pricing data. If OpenAI and Anthropic are forced into further price cuts in response to V4.1 Flash's September 14 routing change, the margin story deteriorates. If they hold prices and retain enterprise customers, the moat holds.
Outlook: Three Horizons, Three Different Answers
Short term, sentiment rules. The September 10-11 moves reflect a repricing of the risk that compute demand could disappoint, and that repricing can overshoot in both directions - as the January 2025 reversal demonstrated. Traders will watch the next hyperscaler earnings calls and any HBM pricing data for confirmation.
Medium term, fundamentals take over. The base case is a split outcome: chip volumes keep growing as agentic AI adoption broadens, but revenue per unit of compute falls, compressing margins at the model layer faster than at the infrastructure layer. The upside case is that usage growth outruns efficiency gains, restoring pricing power to the labs and demand visibility to the suppliers. The downside case is that enterprise budgets do not expand with falling prices, leaving both layers fighting for a smaller pool of dollars.
Long term, the structural question is whether AI inference becomes a commodity. DeepSeek's MIT-licensed open weights mean any company can build on its models without asking permission, and that is the mechanism by which pricing power migrates down the stack. If open-weight models remain within striking distance of the frontier, the model layer looks less like a moat and more like a toll road with competing lanes. The winners in that world are the companies that make inference cheaper and more sovereign - and the investors who priced the AI trade for permanent scarcity may have priced the wrong asset.
The bottom line: DeepSeek's new model is not a market-shattering surprise like R1 was in January 2025 - it is something more dangerous for the AI trade. It is a predictable, repeatable demonstration that frontier capability keeps getting cheaper, and this time the market is being asked to price that reality into valuations built on scarcity. The chip suppliers still have volume on their side, but the pricing-power premium is no longer theirs to keep.
Explore more exclusive insights at nextfin.ai.
