"Reproduce the 'DeepSeek Moment'? Wall Street says: Kimi K3 actually strengthens computing power demand."

CN
13 hours ago

Original author: Long Yue

Original source: Wall Street Insights

The market panicked over Kimi K3 as a "DeepSeek moment 2.0," but this time, Wall Street's judgment is completely different.

Late at night on July 16, Moonlight released Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis intelligence index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More importantly, on the Frontend Code Arena programming leaderboard created by the University of California, Berkeley, K3 topped with a score of 1679, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to surpass all overseas closed-source models on an authoritative programming leaderboard.

On July 17, the US stock semiconductor sector saw a significant drop. The market's reflex is understandable— in early 2025, the release of DeepSeek R1 triggered a sharp drop in computing stocks, based on the logic: If Chinese models are getting stronger, do American AI companies still need to spend so much on computing power? If Chinese models can approach leading capabilities at a lower cost, will the demand for Nvidia, HBM, servers, and network equipment be re-evaluated?

However, according to news from Chasing Wind Trading Desk, the latest research reports from investment banks including UBS, Nomura, Bank of America Merrill Lynch, and Citigroup believe: Kimi K3 is not the end of computing power demand, but an accelerator.

Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 made the market see "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context, always-on inference, native multimodal capabilities, and MoE architecture, these features are not a light asset story. They will elevate the pressure on inference, memory, network, and storage together.

How strong is Kimi K3?

Kimi K3 was released by Moonlight on July 16, 2026, with complete model weights scheduled to be released on July 27. It is an open-source large model with 2.8 trillion parameters, regarded by multiple institutions as the largest-scale open-source weights LLM currently available.

The core configuration includes three points:

First, a 1M token context window. The model can handle longer texts, larger code bases, more complex corporate profiles, and research tasks.

Second, always-on inference. Not just simple Q&A, but aimed at long-chain reasoning and agent tasks.

Third, native visual capabilities. K3 processes not just text but also targets multimodal tasks such as video, images, game development, front-end design, CAD, etc.

In terms of architecture, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE component activates 16 experts for each token out of 896 experts. Moonlight states that compared to Kimi K2, overall scaling efficiency has improved by about 2.5 times.

This explains why K3 is not "a cheaper K2." According to the pricing compiled by Nomura, K3's input cost is $3 per million tokens, cache hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis's metrics, K3's cost per task is about $0.94. This price is lower than Claude Fable 5's approximately $2.75 and Claude Opus 4.8's approximately $1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32–$0.47 and far above DeepSeek V4 Pro's $0.04.

Therefore, K3's positioning is not the lowest price, but approaching the capabilities of leading models at a lower price.

Four investment banks intensively clarify: This is not demand weakening

Regarding the market's concerns about the "DeepSeek moment," Nomura Securities' Asia Pacific technology team analyst Duan Bing wrote in a report: "We believe the competition and innovation in the global large model market will not stop. As we get closer to General Artificial Intelligence (AGI), the application of generative AI on both consumer and enterprise sides will continue to expand. Leading AI labs and hyperscale cloud platform companies are likely to continue investing at this stage to maintain competitive positions— as scale economies continue, we interpret this competition as a positive for the AI infrastructure value chain."

Citigroup semiconductor analyst Peter Lee titled his July 19 report directly as "Another Jevons Paradox." What does Jevons Paradox mean? Simply put: increased efficiency of coal steam engines leads to greater coal consumption because more people can afford it and more scenarios can use it. AI models are the same— when high-quality models become cheaper, developers and companies will deploy more applications, process more tokens, and ultimately the power consumption will rise instead.

Peter Lee believes that even if K3 is widely used, the demand for general memory such as server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other leading models, but the KV cache occupation will expand with context growth, thereby increasing pressure on memory rather than decreasing it.

Bank of America Merrill Lynch semiconductor analyst Vivek Arya stated more directly in a July 17 report. He believes that the response of leading American AI labs is "not less computing power, but more." If Chinese open-source models continue to approach, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also mentioned an easily overlooked background: media reports indicate that Google's Gemini 3.5 Pro has been delayed by several months compared to plans, and programming performance has not met internal targets, making it "increasingly difficult to defend" their leading position.

The UBS analyst Timo Arcuri team pointed out in a report on July 20 that while K3 and DeepSeek R1 do have parallels, K3 is more about scale— it is the world's largest open-source model, with 2.8 trillion parameters and a 1 million token context window. Analysts emphasized that open-source models generally consume more memory than closed-source leading models because longer context windows entail continuous growth in KV cache demand, thus making deployment of open-source models even more reliant on HBM and storage.

Who truly benefits in this competition?

Storage: the most direct benefiting sector. UBS estimates that the cumulative free cash flow (FCF) of the storage and memory sectors is expected to reach about 30% of market value by 2028, the highest among all sub-sectors—Micron (MU) alone accounts for 47%. Citigroup and Nomura both maintain buy ratings on Samsung Electronics, citing the global memory market being in an extremely tight supply situation. Citigroup analyst Peter Lee noted that Kimi K3's inference side memory demands are no less than those of other leading models, and the expansion of KV cache volume will directly drive demand for server DDR5 and enterprise-grade solid-state drives (eSSD). He particularly pointed out that Kimi K3's large-scale deployment requires "super node" cluster configurations with over 64 GPUs.

Computing power infrastructure: TSMC and Nvidia will benefit the most. Whether it's the effectiveness of the training side scale economies or the growth of inference side token demand, it ultimately leads to increased demand for more advanced process chips. Nomura reiterated buy ratings for TSMC, ASE, MediaTek, and others. Nvidia has publicly stated that the inference of modern MoE (Mixture of Experts architecture) models, in terms of performance-to-power ratio on GB300 NVL72, has improved by as much as 25 times compared to the previous generation Hopper architecture. Models like K3 are naturally beneficiaries of Nvidia's latest hardware.

Network: the supernode trend creates structural opportunities. Kimi K3 requires super node clusters, and the limitation of domestic computing power in China due to high-end chip export controls necessitates reliance on super node architecture to make up for performance gaps on single cards, which drives demand from network layer vendors like optical modules and optical chips. Nomura is optimistic about Zhongji Xuchuang and Suzhou Xuchuang.

Cloud platforms: benefit from ecological aggregation effects. Cloud platforms that host various leading open-source models have stronger bargaining power, not relying on a single closed-source model supplier. Nomura is optimistic about Alibaba (BABA) as the core of China's AI cloud ecosystem, as well as data center operators like GDS and VNET.

How fast is the global penetration of Chinese AI models?

This may be the most easily underestimated data point in the entire narrative.

According to statistics from the open API gateway OpenRouter, the token usage of Chinese AI models now accounts for over 45% of global developer traffic, up from less than 2% a year ago. Bank of America Merrill Lynch's data corroborates the acceleration of overall AI penetration: currently about 55% of American enterprises have subscribed to AI models, platforms, or tools, with Anthropic’s enterprise adoption rate at 42% and OpenAI at 40%. Top AI consumers (the top 1% of enterprise users) have reached $4,833 in AI spending per employee per month.

The market is differentiating. On one hand, there are Chinese open-source models represented by DeepSeek and Kimi K3, covering the budget and mid-to-high-end cost-performance markets; on the other hand, top American leading models are focusing on more complex workloads (such as scientific computing), maintaining technological and pricing premiums. Nomura's judgment is that leading large model players on both the US and Chinese sides will benefit— provided they can both continue to stay at the forefront of the technological curve.

K3 truly changes the pace of competition, not just a single company's story

After K3's release, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful, but it should not stop at the level of "Are we going to invest in AI hardware again?"

DeepSeek led the market to reassess training efficiency. K3 allows the market to see another thing: open-source models can also push scale, long context, agents, and multimodality into the forefront.

This will compel American leading labs to continue investing, and it will also allow Chinese models to continue expanding in the global developer ecosystem. Closed-source top models retain technology and pricing premiums, while open-source models cover more price ranges and deployment scenarios, with cloud vendors providing model distribution and enterprise implementation, and the hardware supply chain bearing the training and inference pressure.

In the short term, trading may fluctuate due to "DeepSeek memories." In the medium term, as long as token usage continues to grow, long context and agents continue to spread, computing power, HBM, storage, network, and IDC will still be unavoidable cost items.

This is also the reason multiple institutions have given similar conclusions after K3: stronger open-source models are not the end of AI infrastructure demand, but may instead be the entry point for the next round of demand dispersion.

However, Bank of America Merrill Lynch also clearly left a tail risk: "If the speed of efficiency gains exceeds the growth of workload, we may see some pullback in infrastructure construction." In other words, if models become cheaper, but usage does not significantly expand in tandem, the logic for increased computing power demand will be discounted.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink