Buy the Dip? Why China’s Kimi Model Is Actually Great News for Nvidia
Billy Duberstein, The Motley Fool
Wed, July 22, 2026 at 5:05 PM GMT+5:30
5 min read
- NVDA
+2.30%
Shares of Nvidia (NASDAQ: NVDA) and most of the AI-related semiconductor sector sold off last week after Moonshot, a China-based AI start-up, released its Kimi 3 model
Kimi made waves across the industry, as the open-weights model displayed impressive performance against even the latest frontier models by Anthropic and OpenAI
Missed Nvidia in 2009? This Rare Signal Is Flashing Again. In 2009, a “Double Down” signal flashed for a little-known chipmaker called Nvidia. For the first time in years, that same “Total Conviction” signal is flashing for a company 1/100th the size of Nvidia. Continue »
But the knee-jerk reactions to Kimi 3 seem like an echo of the DeepSeek and TurboQuant sell-offs of early 2025 and 2026, respectively. In both cases, innovations that made AI much more efficient didn’t derail the AI build-out; in fact, one could argue they accelerated it by lowering adoption costs
While these past cases aren’t perfect mirrors of Kimi 3, here’s why Nvidia investors shouldn’t panic over this new model
Why Kimi sent a shudder through U.S. AI stocks
Although Moonshot and other Chinese AI labs may have smuggled in some Nvidia chips illegally, Moonshot likely doesn’t have access to nearly as many Nvidia chips for model training as the leading U.S. labs. There is also some uncertainty about whether Moonshot merely “distilled” a leading LLM from either Anthropic or OpenAI, essentially copying the weights from the U.S. labs
Either way, Kimi 3 appears to have been trained at a small fraction of the cost of leading U.S. models, leading to panic over whether the U.S. giants should and will keep spending on high-end, very expensive Nvidia GPUs
Another reason why Kimi may have spurred a sell-off in Nvidia and AI memory stocks is that it displayed a novel innovation called Kimi Delta Attention (KDA). This architecture enables the model to selectively read prior tokens to process new ones, rather than reading all prior tokens. The result is a 75% decline in KV cache, essentially an AI’s short-term memory required to run the model, and a sixfold increase in speed. That means the model requires less memory and processing power, all things being equal.
Kimi doesn’t lower inference requirements as much as feared
Regardless of how Kimi was trained, if consumers and enterprises want to use it, the model has to run. And while KDA certainly makes more efficient use of KV cache, other architectural features make it somewhat compute-intensive, requiring high-end hardware such as the latest Nvidia racks

