Every published Top AI Stories item tagged with Moonshot AI, newest first.
The inference company Wafer published benchmarks on July 31 putting Moonshot's Kimi K3, at 2.8 trillion parameters, at 952 tokens per second per node on AMD's MI355X against 1,568 on Nvidia's B300. The AMD part is slower in raw throughput but roughly 2.4-times cheaper per GPU, which works out to 48 tokens per second per dollar versus 33. Its 288 gigabytes of memory per GPU also lets the model fit in a single node where the Nvidia configuration needs two.
Moonshot AI has published downloadable weights for Kimi K3, a 2.8 trillion parameter mixture-of-experts model that activates 104 billion parameters per token and accepts a 1 million token context window. It is the first open-weight release in the 3-trillion-parameter class, and it ships with native vision and a custom license permitting commercial use with conditions. Running it is the hard part — the full checkpoint spans 96 shards and well over a terabyte, so most developers will meet it through hosted endpoints rather than their own hardware.
The UK AI Security Institute and the US Center for AI Standards and Innovation published a joint assessment of Moonshot AI's Kimi K3, finding it a measurably weaker offensive-cyber tool than frontier American models but ahead of the leading open-weight alternative. On one simulated attack range Kimi K3 reached step 17 of 32 on average, against 28.5 for top US models, and it failed to develop a working arbitrary-code-execution exploit on any of 41 test cases. The institutes also warned that its safeguards did not stop it from attempting exploit development — a caution as the model stays at the center of the US-China open-weight debate.
Nearly 200 venture-backed startups, organized by the newly formed Little Tech Association, sent a letter urging the Trump administration not to ban US access to Chinese open-weight AI models. Signatories including Y Combinator and Proton argue that cheap open models from Moonshot AI and Alibaba are a lifeline for small companies, and that a broad prohibition would raise costs and hand even more power to a few dominant US labs. They asked for "targeted safeguards" rather than a blanket cutoff.
Feeding that same debate, AI researchers pushed back hard on the White House claim that Moonshot's Kimi K3 reached the frontier by secretly distilling Anthropic's Claude Fable 5. Skeptics point to the timeline — Fable 5 went public on July 1 and Kimi K3 shipped just 14 days later — and note that no logs or forensic evidence have been released. Others argue model outputs are not copyrightable in the first place, undercutting the administration's framing of distillation as technology theft.
The White House's technology chief, Michael Kratsios, publicly accused Chinese lab Moonshot AI of running large-scale distillation against Anthropic's Fable model to help build its Kimi K3 system — and of training on export-controlled Nvidia GB300 servers accessed through Thailand. Treasury Secretary Scott Bessent said sanctions and Entity List designations are "on the table." Some researchers are skeptical distillation alone could explain Kimi K3, noting Anthropic only released Fable publicly on July 1.
US Treasury Secretary Scott Bessent said Washington will examine leading Chinese open-weight models — most recently Moonshot AI's Kimi K3 — for signs they were "distilled" from American systems, and warned the US could impose sanctions if it finds IP theft. Distillation, which transfers a large model's capabilities into a smaller one, is a common industry technique that American labs use too, and Microsoft's Satya Nadella has called the theft framing "ironic." The move marks a sharp escalation in the US-China frontier-model race.
In this week's Stratechery, Ben Thompson argues that Western alarm over cheap Chinese open-weight models — Moonshot's Kimi K3 and Alibaba's forthcoming Qwen3.8 Max — misreads the economics. What matters, he writes, is not the sticker price per token but the total cost of a correct answer, since different models burn very different token volumes to solve the same problem. As intelligence becomes a commodity, he contends, US frontier labs with better cost structures still win — and he urges loosening rules that push American security teams toward Chinese models.
China's Moonshot AI released Kimi Work, a downloadable desktop agent for Windows and Mac that reads local files, automates a browser, runs scheduled tasks around the clock, and coordinates a swarm of specialized sub-agents to build slide decks and spreadsheets. It ships with built-in global market data and can run Python and shell scripts. The launch pushes Moonshot beyond its chatbot roots into the same agentic-desktop territory as OpenAI and Anthropic's newest cowork products.
Moonshot AI released Kimi K3, a mixture-of-experts (MoE) model with 2.8 trillion total parameters, a one-million-token context window and native vision, activating 16 of 896 experts per token. It beats Claude Opus 4.8 on several coding benchmarks — 88.3 versus 84.6 on Terminal Bench — while Moonshot concedes it still trails Claude Fable 5 and GPT-5.6 Sol overall, and names user experience as the remaining gap. Full weights arrive by July 27, and the company is reportedly raising at a $31.5 billion valuation, up from $20 billion in May.
GitHub started rolling out Moonshot AI's Kimi K2.7 Code — an open-weight model from the Chinese lab — inside GitHub Copilot for Pro, Pro+, and Max subscribers, with business and enterprise plans to follow. GitHub hosts the model on Microsoft Azure and bills it by usage. It is a notable sign of Chinese open models reaching the West's most widely used coding assistant, though enterprise admins must switch it on manually before their teams can pick it.
China's Moonshot AI released Kimi K2.7 Code, a one-trillion-parameter open-weights coding model that the lab says cuts reasoning-token use by about 30 percent while topping its previous K2.6 release on internal coding benchmarks. Moonshot published the full model weights to Hugging Face the same day and priced it well below Western flagships. But practitioners quoted by VentureBeat cautioned that the headline benchmark gains have been hard to reproduce in real-world use — a recurring pattern as open-source labs race to claim coding leadership.
China's Moonshot AI released Kimi Work, a downloadable desktop agent for macOS and Windows that runs on its open-weight Kimi K2.6 model. Unlike a web chatbot, it works directly with local files and drives the user's logged-in browser through an extension, and it can fan a task out across as many as 300 parallel sub-agents. The launch lands a day after Moonshot's reported $2 billion raise, underscoring how fast Chinese labs are shipping agentic products.
Moonshot AI, the maker of the Kimi chatbot, is in early talks to raise as much as $2 billion at a $30 billion valuation — a sevenfold jump from the roughly $4 billion the startup was worth in December. The round, its third in six months, comes as Chinese labs race to keep pace with each other and with Western rivals. Moonshot's annualized revenue passed $200 million in April, and its upgraded Kimi K2.6 model ranks among the most-used models on the OpenRouter distribution platform. The pace shows how fast capital is flowing into China's open-model contenders even as US export controls bite.
The EAGLE Team, vLLM, and TorchSpec released EAGLE 3.1 on May 26, a speculative-decoding update that addresses "attention drift" by adding FC normalization after each target hidden state and feeding post-norm states into subsequent decoding steps. On Moonshot's Kimi K2.6 base model with vLLM, the team reports a 2.03-times per-user throughput speedup at concurrency one and a 1.66-times speedup at concurrency 16 on Nvidia GB200 hardware, plus up to twice the acceptance length in long-context scenarios versus EAGLE 3. The update maintains backward compatibility with existing EAGLE 3 checkpoints.
Cursor released Composer 2.5, an updated coding model built on Moonshot AI's open-source Kimi K2.5 checkpoint and trained with 25-times more synthetic tasks than Composer 2, plus a new sharded-Muon distributed-training setup. Cursor claims "substantial improvement in intelligence and behavior" on long-running tasks and complex instruction-following. Pricing is $0.50 per million input tokens and $2.50 per million output, with a fast variant at $3 and $15. The first week ships with double usage included.
Beijing-based Moonshot AI closed $2 billion led by Meituan's Long-Z Investments arm, with Tsinghua Capital, China Mobile, and CPE Yuanfeng participating. The round nearly doubles Moonshot's valuation from earlier this year and tracks $200 million annualized recurring revenue, as Kimi K2.6 climbed to the second-most-used model on OpenRouter. Coming a day after DeepSeek's reported $45 billion talks, the round signals open-weights labs out of China are emerging as the primary cost-pressure on US frontier vendors.