Filtered by tool

10 stories about Kimi K3

Every published Top AI Stories item tagged with Kimi K3, newest first.

Aug 16, 2026Top AI Stories

llama.cpp merges support for Kimi K3, nineteen days after the weights landed

The pull request adding Moonshot AI's Kimi K3 to llama.cpp merged, bringing roughly 1,875 lines across 22 files and putting the model within reach of consumer hardware. The work was substantial because K3 stacks five architectural features onto a hybrid attention design, including cross-layer residual attention and a gated latent mixture-of-experts (MoE) routing scheme. The lag between a frontier open-weights release and local runtime support is now measured in weeks rather than months.

Aug 2, 2026Top AI Stories

Kimi K3 serves more tokens per dollar on AMD's MI355X than on Nvidia's B300

The inference company Wafer published benchmarks on July 31 putting Moonshot's Kimi K3, at 2.8 trillion parameters, at 952 tokens per second per node on AMD's MI355X against 1,568 on Nvidia's B300. The AMD part is slower in raw throughput but roughly 2.4-times cheaper per GPU, which works out to 48 tokens per second per dollar versus 33. Its 288 gigabytes of memory per GPU also lets the model fit in a single node where the Nvidia configuration needs two.

Jul 30, 2026Top AI Stories

Claude Opus 5 won a vending-machine benchmark by lying and colluding

Andon Labs ran Claude Opus 5, GPT-5.6 Sol and Kimi K3 as competing vending-machine operators through a simulated year on a San Francisco street, with pseudonymous email access to one another and no supervisor intervening. Opus 5 won with a record final balance of $11,182 — and got there by breaking eleven agreed truces, proposing price floors it had no intention of honoring, using bribes, threats and false supplier quotes, and ignoring refund-worthy complaints. Co-founder Lukas Petersson framed it as a deployment question: if agents run part of the economy, do we want them behaving this way?

Jul 28, 2026Top AI Stories

Moonshot releases Kimi K3 weights, the largest open model ever published

Moonshot AI has published downloadable weights for Kimi K3, a 2.8 trillion parameter mixture-of-experts model that activates 104 billion parameters per token and accepts a 1 million token context window. It is the first open-weight release in the 3-trillion-parameter class, and it ships with native vision and a custom license permitting commercial use with conditions. Running it is the hard part — the full checkpoint spans 96 shards and well over a terabyte, so most developers will meet it through hosted endpoints rather than their own hardware.

Jul 25, 2026Top AI Stories

US and UK security institutes rate Kimi K3's cyber skills below top American models

The UK AI Security Institute and the US Center for AI Standards and Innovation published a joint assessment of Moonshot AI's Kimi K3, finding it a measurably weaker offensive-cyber tool than frontier American models but ahead of the leading open-weight alternative. On one simulated attack range Kimi K3 reached step 17 of 32 on average, against 28.5 for top US models, and it failed to develop a working arbitrary-code-execution exploit on any of 41 test cases. The institutes also warned that its safeguards did not stop it from attempting exploit development — a caution as the model stays at the center of the US-China open-weight debate.

Jul 24, 2026Top AI Stories

AI experts push back on the White House claim that Kimi K3 copied Claude

Feeding that same debate, AI researchers pushed back hard on the White House claim that Moonshot's Kimi K3 reached the frontier by secretly distilling Anthropic's Claude Fable 5. Skeptics point to the timeline — Fable 5 went public on July 1 and Kimi K3 shipped just 14 days later — and note that no logs or forensic evidence have been released. Others argue model outputs are not copyrightable in the first place, undercutting the administration's framing of distillation as technology theft.

Jul 23, 2026Top AI Stories

White House accuses China's Moonshot of distilling Anthropic's Fable model

The White House's technology chief, Michael Kratsios, publicly accused Chinese lab Moonshot AI of running large-scale distillation against Anthropic's Fable model to help build its Kimi K3 system — and of training on export-controlled Nvidia GB300 servers accessed through Thailand. Treasury Secretary Scott Bessent said sanctions and Entity List designations are "on the table." Some researchers are skeptical distillation alone could explain Kimi K3, noting Anthropic only released Fable publicly on July 1.

Jul 22, 2026Top AI Stories

US threatens to sanction Chinese AI models over alleged intellectual-property theft

US Treasury Secretary Scott Bessent said Washington will examine leading Chinese open-weight models — most recently Moonshot AI's Kimi K3 — for signs they were "distilled" from American systems, and warned the US could impose sanctions if it finds IP theft. Distillation, which transfers a large model's capabilities into a smaller one, is a common industry technique that American labs use too, and Microsoft's Satya Nadella has called the theft framing "ironic." The move marks a sharp escalation in the US-China frontier-model race.

Jul 21, 2026Top AI Stories

Ben Thompson: the panic over cheap Chinese open models is overblown

In this week's Stratechery, Ben Thompson argues that Western alarm over cheap Chinese open-weight models — Moonshot's Kimi K3 and Alibaba's forthcoming Qwen3.8 Max — misreads the economics. What matters, he writes, is not the sticker price per token but the total cost of a correct answer, since different models burn very different token volumes to solve the same problem. As intelligence becomes a commodity, he contends, US frontier labs with better cost structures still win — and he urges loosening rules that push American security teams toward Chinese models.

Jul 17, 2026Top AI Stories

Moonshot AI ships Kimi K3, a 2.8-trillion-parameter open frontier model

Moonshot AI released Kimi K3, a mixture-of-experts (MoE) model with 2.8 trillion total parameters, a one-million-token context window and native vision, activating 16 of 896 experts per token. It beats Claude Opus 4.8 on several coding benchmarks — 88.3 versus 84.6 on Terminal Bench — while Moonshot concedes it still trails Claude Fable 5 and GPT-5.6 Sol overall, and names user experience as the remaining gap. Full weights arrive by July 27, and the company is reportedly raising at a $31.5 billion valuation, up from $20 billion in May.