NewMoneyMoves
PerspectivesResearchTerminalsWeekend RWABoard
THEME
EnergyGlobal ThemeFront-runBioRoboticsMobilityInfra
››Theme Deep Dive
GLOBAL THEME

The Great Open-Weight Migration — the Four-Layer Remapping of AI Margins After K3

Closed labs, serving providers, hyperscalers, neoclouds — where China's giant open-weight model evaporates AI margins, and where it moves them

HHaelangdal·Founder AnalystJuly 20, 202659 min readTheme Deep Dive
Bottom Line

K3 is not DeepSeek's second coming but its mirror image. The compute to build was compressed, the compute to serve got bigger, and margins are migrating from models to infrastructure rather than vanishing. The benefit ranking runs hardware (memory included) > AWS and Google > spot neoclouds and heavyweight serving, with pressure concentrated on second-tier closed models and long-tail serving. For memory it is neutral-or-better as long as capex holds — K3's 288GB-per-GPU serving requirement is, if anything, an HBM mix-upgrade driver.

NewMoneyMoves

Global Investment through Themes

Content

ThemesIdeasReportsEarnings

Explore

GuideETFsPerspectivesTelegramNVIDIA PortfolioWeekly ReviewCredit & Leverage

Community

Daily News

Legal

Privacy PolicyTerms of ServiceDisclaimer

© 2026 NewMoneyMoves. All rights reserved.

Reader's Brief — 30-second TL;DR

Advanced
Why Now

7/16: Moonshot AI releases Kimi K3, the largest open-weight model ever (2.8T parameters, $3/$15) → SOX -20.2% from peak into a technical bear market, global semiconductor market cap -$3.3T → the same day, serving provider Fireworks raises $1.5B at a $17.5B valuation → most of the correction predates K3 (the 7/1 Meta Compute report, a hawkish Fed, Iran) — a multi-cause structure.

Winners ?? Losers

Beneficiaries — hardware (NVIDIA, SK Hynix, memory), model-neutral Amazon AWS, vertically integrated Google, spot neoclouds, heavyweight serving and specialty hardware. Neutral — contract-based neoclouds (the CoreWeave type), Microsoft (OpenAI-derivative exposure). Pressured — second-tier closed models (the Opus 4.8 and GPT-5.5 price band), the generic small-model serving long tail, China's second-tier labs (Zhipu -28.4%, MiniMax -15.6%).

Watch For

7/22 Alphabet capex guidance → 7/27 K3 weights release, third-party quotes, CXMT Shanghai listing → week of 7/29 Microsoft, Meta, Amazon and SK Hynix earnings → August: Korea's leveraged ETF margin rules take effect → Sep–Q4 OpenAI IPO window (2027 delay under review) → whether Bedrock and Vertex host K3.

0%
Reading depth

Summary — Margins Do Not Vanish. They Migrate

On July 16, 2026, Moonshot AI released Kimi K3, the largest open-weight model in history at 2.8 trillion total parameters. By the company's own benchmarks it trails the top frontier tier, but it beat the second-tier frontier models on coding and agentic tasks, and the full are scheduled for release on July 27. The market immediately summoned the memory January 2025's DeepSeek moment. The Philadelphia Semiconductor Index (SOX) fell 20.2% from its June 22 all-time high into a technical bear market, and roughly $3.3 trillion of global semiconductor market capitalization has evaporated since June 22.

Yet the two events resemble each other only on the surface. Structurally they are near opposites. DeepSeek was a story of "reaching the frontier small and cheap," which spawned a compute-glut narrative. K3 is a story of "building big and selling frontier-adjacent performance at US frontier-adjacent prices." Even with FP4 quantization, K3 does not fit on a single NVIDIA DGX B200 system — it demands GB300 NVL72-class systems with 288GB of memory GPU. The scaling laws are not broken. The price tag of the compute required for scaling is simply falling faster than US labs assumed.

What K3 actually hits is not the semiconductor supply chain but the margin allocation structure of the AI value chain. The question "what happens to cloud provider margins" has no answer as a single lump. Split the stack into four layers and the sign of the shock differs by layer. Layer 1 (closed labs) faces a direct hit to the frontier-gap premium while agent lock-in (switching costs that bind customers) and the regulatory-trust premium remain defended — an asymmetric structure. Layer 2 (serving specialists) is the front line of margin compression, yet capital is flowing in the opposite direction — Fireworks closed a $1.5 billion round at a $17.5 billion valuation on the very day K3 was announced. For Layer 3 (hyperscalers), open weights are not a threat but a rerun of the playbook that monetized Linux, with the benefit ranking running Amazon AWS, Google, then Microsoft. Layer 4 (neoclouds) must be split into contract-based and spot-based operators: the benefit of self-hosting demand lands on spot players, while credit-transmission risk concentrates in contract players.

In sum, the direction of the margin migration is top-down — from models to infrastructure. The biggest beneficiary is the hardware layer, memory included; next come AWS with its model-neutral resale structure, vertically integrated Google, spot-based neoclouds, and specialty-hardware and heavyweight serving providers. For memory, this reshuffle is neutral to positive. Tokens run on DRAM no matter which layer serves them, and K3's own hardware requirements are the proof. There is exactly one malignant path: Layer 1 margin compression feeding through an OpenAI IPO pricing disappointment into capital-market doubt about AI monetization, and from there into hyperscaler capex cuts — a four-step . Observation of the first domino begins with Alphabet's earnings on July 22, runs through the big-tech earnings of the week of July 29, and extends to OpenAI's listing timeline.

For Korea's value chain, the short version: the open-weight migration does not damage the fundamentals of the memory leaders. SK Hynix closing higher the day after K3's release, while US semis slid — though it overlapped with an oversold bounce after steep prior losses and cannot be read conclusively — can plausibly be read as the market starting to price the proposition that "an open-weight victory is not an HBM defeat." K3's 288GB-per-GPU serving requirement is, if anything, an HBM mix-upgrade driver. Margins do not vanish. They change address.

The Event — K3 and the Second DeepSeek Moment

Key Points
  • —K3 is not DeepSeek's second coming but its mirror image — the compute to build was compressed, while the compute to serve got bigger.

Specs, architecture, price

Moonshot AI released Kimi K3 on July 16, US Eastern time. It is a mixture-of-experts (MoE) architecture with 2.8 trillion total parameters; 16 of 896 experts activate per token, so roughly 50 billion parameters work per forward pass. It introduces the in-house KDA hybrid linear attention and attention-residual techniques, supports native visual understanding and a 1-million-token context window. It is a roughly 2.8x jump from K2's 1 trillion parameters, far above DeepSeek V4 Pro (1.6T) and Zhipu's GLM 5 family (744B) — the largest open-weight model in the world. The full weights are due July 27, and the license is flagged to be the same modified MIT as its predecessor (with brand-attribution obligations for very large users).

Total parameters of major Chinese open-weight models — from GLM 5 (0.744T), Kimi K2 (1T) and DeepSeek V4 Pro (1.6T) to Kimi K3's jump to 2.8T. Source: Moonshot AI, company disclosures
Total parameters of major Chinese open-weight models — from GLM 5 (0.744T), Kimi K2 (1T) and DeepSeek V4 Pro (1.6T) to Kimi K3's jump to 2.8T. Source: Moonshot AI, company disclosures

<Chart 01> Total parameters of major Chinese open-weight models — from GLM 5 (0.744T), Kimi K2 (1T) and DeepSeek V4 Pro (1.6T) to Kimi K3's jump to 2.8T. Source: Moonshot AI, company disclosures

The performance position is clear. By Moonshot's own evaluation, K3's overall performance trails the top frontier — Claude Fable 5 and GPT-5.6 Sol — but it consistently beat Claude Opus 4.8 and GPT-5.5 on programming, visual understanding, and long-horizon tasks. Within a day of release it took the #1 spot on the frontend-coding leaderboard of Arena, built by UC Berkeley researchers, and ranked ninth in the world on the overall text leaderboard as of the immediate post-release snapshot. Arena's CEO Angelopoulos called it possibly the biggest single release of the year. Output-token efficiency also improved — completing the same evaluation set with about 132 million tokens, 21% fewer than K2.6's roughly 166 million, while scoring 13 points higher on the intelligence index.

Price is what defines the character of this event. API pricing is $3 per million input tokens and $15 per million output tokens — the highest in Chinese AI lab history, roughly a 3x input and 4x output increase over K2.6. A 90% discount drops input to $0.30 on cache hits, so effective rates fall sharply for long-document and long-running agent workloads, but at sticker price this is US second-frontier territory. In Artificial Analysis's independent evaluation, K3's average cost per task is $0.94 — effectively the same bracket as GPT-5.6 Sol ($1.04) and about half of Opus 4.8 ($1.80), but 3x GLM-5.2 ($0.32) and more than 20x DeepSeek V4 Pro ($0.04). This is why it reads as the end of the ultra-cheap Chinese AI era.

Average evaluation cost per task by model — K3 ($0.94) sits in the same bracket as GPT-5.6 Sol ($1.04), about half of Opus 4.8 ($1.80). Source: Artificial Analysis
Average evaluation cost per task by model — K3 ($0.94) sits in the same bracket as GPT-5.6 Sol ($1.04), about half of Opus 4.8 ($1.80). Source: Artificial Analysis

<Chart 02> Average evaluation cost per task by model — K3 ($0.94) sits in the same bracket as GPT-5.6 Sol ($1.04), about half of Opus 4.8 ($1.80). Source: Artificial Analysis

The supplier, Moonshot AI, is a Beijing-based startup of roughly 300 employees. It was founded by Yang Zhilin, a Carnegie Mellon PhD who worked with Google Brain and Meta research. The consumer Kimi chatbot passed $200 million in annualized revenue (ARR) as of April, and in May the company raised $2 billion at a valuation above $20 billion. Consumer subscriptions sell in four paid tiers from $19 to $199 a month.

January 2025 vs July 2026: two shocks, two structures

Replaying the DeepSeek moment sharpens the contrast. The January 2025 shock started from a number: V3's $5.57 million training cost. The claim that a model activating only 37B of 671 billion (671B) total parameters matched then-frontier performance spawned the reading that "reaching the frontier costs one-hundredth of what US labs assume," which fed straight into a "compute will be left over" oversupply narrative and NVIDIA's 17% single-session crash. The path afterward is well known. Reasoning-model usage exploded, compute demand surged instead, and the stock recovered to new highs — a textbook demonstration of the , where efficiency gains feed back into demand expansion.

K3's narrative differs from the starting point. First, the direction is reversed. DeepSeek was "small and cheap"; K3 is "big, at frontier prices." Second, the hardware implication is reversed. Per SemiAnalysis, even with FP4 quantization K3 does not fit a single DGX B200, requiring GB300 NVL72 or B300-class systems with 288GB per GPU. Where DeepSeek bred the misreading that "less compute is needed," K3 is by construction "a model that needs top-spec systems to serve." Third, the intensity of the market reaction differs. Unlike DeepSeek's 17% single-day NVIDIA crash, on July 17 NVIDIA fell 3% intraday and closed down 2.2%, Micron closed roughly flat (-0.5%), and SK Hynix actually closed higher. The market remembers the lesson of DeepSeek — that the panic was overdone and turned out to be a buying opportunity.

ItemDeepSeek moment (Jan 2025)K3 (Jul 2026)
Core narrativeReaching the frontier small and cheap (V3 training cost $5.57M)The largest model ever, open-weight, at frontier-adjacent prices
Model size671 billion (671B) total / 37B active2.8T total / ~50B active (16 of 896 experts)
Price positionExtreme discount vs USSame bracket as GPT-5.6 Sol ($0.94 vs $1.04 per task)
Hardware implicationSpawned a "compute glut" narrativeDoes not fit DGX B200 even at FP4; needs 288GB/GPU class
Market reactionNVIDIA -17% in one sessionNVIDIA -2.2%, much of the intraday loss recovered
Nature of the correctionBolt-from-the-blue single shockOne trigger atop a pre-existing correction (-20% from peak)
AftermathJevons paradox, compute demand surge, new highsVerdict due at the 7/27 weights release and late-July big-tech earnings
Item
Core narrative
DeepSeek moment (Jan 2025)
Reaching the frontier small and cheap (V3 training cost $5.57M)
K3 (Jul 2026)
The largest model ever, open-weight, at frontier-adjacent prices
Item
Model size
DeepSeek moment (Jan 2025)
671 billion (671B) total / 37B active
K3 (Jul 2026)
2.8T total / ~50B active (16 of 896 experts)
Item
Price position
DeepSeek moment (Jan 2025)
Extreme discount vs US
K3 (Jul 2026)
Same bracket as GPT-5.6 Sol ($0.94 vs $1.04 per task)
Item
Hardware implication
DeepSeek moment (Jan 2025)
Spawned a "compute glut" narrative
K3 (Jul 2026)
Does not fit DGX B200 even at FP4; needs 288GB/GPU class
Item
Market reaction
DeepSeek moment (Jan 2025)
NVIDIA -17% in one session
K3 (Jul 2026)
NVIDIA -2.2%, much of the intraday loss recovered
Item
Nature of the correction
DeepSeek moment (Jan 2025)
Bolt-from-the-blue single shock
K3 (Jul 2026)
One trigger atop a pre-existing correction (-20% from peak)
Item
Aftermath
DeepSeek moment (Jan 2025)
Jevons paradox, compute demand surge, new highs
K3 (Jul 2026)
Verdict due at the 7/27 weights release and late-July big-tech earnings

K3 is not DeepSeek's second coming but its mirror image — the compute to build was compressed, while the compute to serve got bigger.

The Market's Scorecard — Causality of the Correction and the Industry Spectrum

Measured market reaction

The numbers first. The SOX hit a record close near 14,655 on June 22, then closed down 1.6% on July 17 — 20.2% below the peak, a technical bear market. The weekly loss of roughly 10–11% was the worst since the April 2025 tariff shock, and July's month-to-date decline exceeds 18%. Global semiconductor market cap has shrunk by about $3.3 trillion since June 22. Still, year-to-date the SOX remains up by nearly triple digits, far ahead of the S&P 500's roughly 9% — meaning this correction is a retracement of a 105% rally (March trough to June peak).

Individual moves on July 17 show where the epicenter was. TSMC fell 7% despite reporting a 77% increase in quarterly net income; SoftBank, which trades as an OpenAI proxy, dropped 9%. Hong Kong-listed Chinese AI rivals took the biggest hits — Zhipu -28.4%, MiniMax -15.6%. In the US, NVIDIA briefly ceded the #1 market-cap spot to Apple intraday, and the S&P 500 was down around 1%. Micron and SanDisk were already down around 30% from their peaks for reasons unrelated to K3.

What matters is the causal structure of this correction. K3 was announced on July 16, but most of the SOX decline happened before it. On July 1, Bloomberg's exclusive report that Meta is pursuing its own cloud (Meta Compute) sent CoreWeave down 13.9% and Nebius down 17% that day, and the SOX fell 6.7% the following day. Layer on the Intel 18A-P yield-delay reports, Netflix's earnings disappointment, Iran-war risk-off, and the hawkish surprise from the Fed under new Chair Warsh — 9 of 18 officials projecting a 2026 hike, the median forecast jumping from 3.4% to 3.8% — and K3 landed on top of an already heavy pile. As JPMorgan's Andrew Tyler put it, K3 "added fuel to the fire" of market anxiety; it was a trigger, not the sole cause. The distinction matters for scenario judgment later — even if the K3 narrative is refuted, the correction's other causes remain.

The industry spectrum

Technical reviews are unusually unanimous in their praise. Michiel Bakker, affiliated with MIT and Google DeepMind, called K3 "insanely good" and wrote that the result looks impossible to explain by distillation alone — distillation being the technique of training a smaller model on a larger model's outputs. The weight of that comment is considerable: the dominant Western hypothesis for Chinese labs' competitiveness — free-riding via distillation — does not apply to K3. Penn's Ethan Mollick rated it "the closest model to the frontier" while noting the jaggedness of its abilities.

On the compute debate, SemiAnalysis's self-correction is symbolic. Just a week before K3's release, SemiAnalysis wrote that Chinese labs lacked the compute to truly reach the frontier. After the release, founder Dylan Patel conceded that an extremely talented small team had closed much of the compute gap through RL, architecture, and data research, and noted that Chinese firms can easily rent GPUs abroad, rendering parts of export controls moot. DeepMind's Anika Somaia posed the more fundamental question — the whole Western consensus, from export controls to the hyperscalers' hundred-billion-dollar arms race to the "compute moat" investment thesis, rests on the single assumption that compute determines capability, while Moonshot's training stack is itself an innovation forced by GPU scarcity. Somaia's formulation supplies one of this report's core axes: a talented small lab can compress the compute needed to build a frontier model, but it cannot afford the compute to serve it.

Reactions inside the competition read differently. Dean Ball, a senior strategy executive at OpenAI (head of strategic futures), acknowledged K3 as a very good model while noting that its heavy token consumption makes it unclear whether it is actually cheap to run. An independent test in which a single simple SVG generation consumed 13,241 reasoning tokens at roughly $0.25 per query supports the point. Ball went further, arguing open-weight models are inherently decelerationist — slowing AI investment — and warning that one endpoint of an open-weight-dominated world is "full AI communism," with states providing AI as public infrastructure. It is a statement to be read with his interests in mind — he is a strategist at a company that depends on a closed business model — but his accompanying policy prediction, of governments using "soft " to seed regulatory uncertainty rather than outright bans, is the starting point for the regulatory-moat discussion below.

Among market interpreters, Patrick Moorhead diagnosed an overreaction strikingly similar to the DeepSeek episode, and Bank of America stressed the fleetingness of AI leadership, noting Moonshot proved that leaps remain possible through better training and design even with restricted access to leading-edge chips. Apollo's Torsten Slok articulated the maximum version of the fear this selloff amplified — to the effect that if capable AI approaches free, the roughly $700 billion hyperscalers are spending (2026 capex consensus) may never be recouped, a potential recession trigger. On the other side, tech analyst Tae Kim argued that K3's near-3-trillion size itself shows the scaling laws still hold and that it will increase compute demand. DeepMind's Demis Hassabis left the most honest summary of the entire spectrum — nobody on Earth knows what happens next.

Early independent verification largely ratified Moonshot's self-grading. On the Artificial Analysis Intelligence Index (v4.1), K3 scored 57.1 — third in the world by model family behind Fable 5 (59.9) and GPT-5.6 Sol (58.9), fourth by configuration behind Sol's high-compute setting, and ahead of Opus 4.8 (55.7). It is the first time a Chinese model has reached this position in independent evaluation. It was stronger still on agentic evaluations — an Elo of 1668 on the knowledge-work benchmark GDPval, above Opus 4.8 (1600) and GPT-5.5 (1494), leaving only Fable 5 (1760) ahead, and #1 on SaaS workflow automation. The real-world usage profile has shadows, though — output speed is an unimpressive ~62 tokens per second, token consumption runs high versus peers, and hallucination metrics regressed.

China's strategic pivot: the end of the ultra-cheap era

K3's price tag should be read as a regime-change signal for the Chinese AI ecosystem. China's 2025 model-export strategy was effectively singular: build developer share through extreme discounts versus the US. DeepSeek V4 Pro's $0.04 per task and GLM-5.2's $0.32 are that legacy. K3 broke from the pack and set its price at $0.94 per task — the same bracket as the US second frontier. That Moonshot raised prices 3x on input and 4x on output versus its predecessor and still saw its servers overloaded within a day of release is the first demonstration that once frontier-adjacent performance is secured, a Chinese model can also exercise pricing power.

The strategic layers are splitting too. Moonshot is frontier-priced open weights; Zhipu holds the ultra-cheap volume lane; DeepSeek keeps the extreme-efficiency franchise. Beijing's policy direction runs through this divergence — Xi Jinping elevating open source and open collaboration to the international agenda in his July 17 opening speech at the World AI Conference (WAIC) in Shanghai signals that global ecosystem expansion via open weights is a national-level direction, not merely a corporate strategy. The market's grading is telling — on K3's release day, Hong Kong-listed rivals Zhipu and MiniMax plunged 28.4% and 15.6%. The rise of Chinese open weights destroys the margin narratives of China's own second-tier labs before it touches the US labs. This game is frontier-versus-non-frontier before it is US-versus-China.

The market shock is real, but it was a trigger atop a multi-cause correction — the real structural change is happening not in stock prices but in margin allocation.

Full access requires 🥉 Bronze tier

Sign in with Google — your tier will be checked automatically and access granted if eligible.

Sign in with Google

This site runs on ads — the tier system rewards community contributions.

Comments

This report is provided for informational purposes only and does not constitute a recommendation to buy or sell any financial instrument. Investment decisions should be made based on your own judgment and responsibility. The analysis and opinions contained herein are based on information available at the time of writing and are subject to change.

All Reports
Home
Research
Company
Macro
Theme