Filtered by tool

15 stories about DeepSeek

Every published Top AI Stories item tagged with DeepSeek, newest first.

Sep 10, 2026Top AI Stories

DeepSeek's newest small model retires its own flagship and ships under MIT

DeepSeek released V4.1-Flash, the smallest model in a new architecture family, and the release notes do something unusual: they retire the previous flagship. From noon Beijing time on September 14, every request to the V4-Pro endpoint will be routed to Flash and billed at the Flash price, because DeepSeek says testing showed the smaller model now beats V4-Pro on performance, cost and speed alike. It is a natively multimodal mixture-of-experts model with a 552 billion parameter backbone that activates only 8 billion per token while reading input, supports a one-million-token context, and cuts the persistent key-value cache to 890 bytes per token, roughly a quarter of the previous generation. The weights are public on Hugging Face under a plain MIT license, with no revenue threshold and no excluded territory, which is increasingly the exception among Chinese open-weight releases. Prices across the API came down with it.

Sep 9, 2026Top AI Stories

Three US agencies accuse six Chinese AI labs of industrial-scale distillation

The National Security Agency, the Federal Bureau of Investigation and the Cybersecurity and Infrastructure Security Agency issued a joint advisory saying DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.ai have extracted billions of tokens from Claude, GPT, Gemini and Grok since late 2024, likely with Chinese government awareness. That escalates a July accusation from the White House technology chief, which named Moonshot alone. Distillation, training a smaller model on a larger one's outputs, is a standard technique — but the agencies say it forms the core rather than a supplement of these labs' development. They also reject the number behind the 2025 market shock, DeepSeek's widely cited training cost of 5.6 million dollars. The advisory says it leaves out the cost of the data taken this way. China's foreign ministry calls the claims groundless.

Sep 1, 2026Top AI Stories

DeepSeek releases its first multimodal V4 model under an MIT license

DeepSeek published DeepSeek-V4-Flash-Vision-Exp, the first multimodal model in its V4 family, as ungated weights under an MIT license. It grafts vision modules onto the V4 Flash architecture, and the lab's own table puts it ahead of Anthropic's Opus 4.8 on two multimodal agent benchmarks, Agents' Last Exam and ZeroBench, while trailing on ApexBench and on most text-agent tests. Those numbers are vendor-reported on DeepSeek's own harness, and the repository ships reference inference code rather than a hosted product, with a vLLM recipe targeting a four-GPU GB300 node.

Aug 15, 2026Top AI Stories

DeepSeek will raise API prices by up to 1,100 percent as OpenAI and Anthropic cut theirs

DeepSeek's own pricing page shows new peak and off-peak rates taking effect on August 16: peak output tokens for its V4 Pro model climb from 87 cents per million to three dollars and ninety-six cents, and some cached-input rates rise roughly 1,100 percent. OpenAI and Anthropic are moving the other way, cutting GPT-5.6 Luna by as much as 80 percent and scrapping a planned September increase for Sonnet 5. TechRepublic, citing SiliconData figures reported in the Financial Times, says DoorDash, Airbnb and Coinbase now run Chinese models in production.

Aug 13, 2026Top AI Stories

DeepSeek's V4 Pro leaves preview after about four months

DeepSeek promoted V4 Pro to general availability as the 0813 build, ending a preview that had run since April. The mixture-of-experts model carries a one-million-token context window and lists at roughly 44 cents per million input tokens and 87 cents per million output — around a fifth of what Grok 4.6 charges on the input side, and a reminder that the price floor for capable inference keeps dropping out of China.

Aug 2, 2026Top AI Stories

A hacker wired DeepSeek into an attack framework and let it run on its own

Palo Alto Networks' Unit 42 documented a Zhuhai-based operator who connected DeepSeek to the open-source Hermes Agent framework and drove it through Telegram. After a single command, the model enumerated targets across ten product families, pulled public exploit code from GitHub, ranked vulnerabilities by severity, and ran the exploitation cycles without further human input — compressing what Unit 42 calls hundreds of hours of manual targeting into minutes. Roughly 460 targets were attempted, with confirmed impact limited to data theft from three Citrix NetScaler systems and command execution on eleven Marimo notebooks. Unit 42 notes that OpenAI's provider-side controls refused the same requests and disabled a linked account.

Aug 1, 2026Top AI Stories

DeepSeek ships V4 Flash weights under a plain MIT license

DeepSeek published the official July 31 build of V4 Flash on Hugging Face under an unmodified MIT license, which allows commercial use with no separate agreement — a sharp contrast with Moonshot's Kimi K3, released four days earlier under custom terms that trigger a bilateral deal above set revenue and user thresholds. Artificial Analysis scores it 50, third among open-weights models. It is a sparse mixture-of-experts (MoE) design activating roughly 13 billion parameters per token, at 14 cents per million input tokens and 28 cents per million output.

Jul 19, 2026Top AI Stories

DeepSeek preps a Shanghai IPO and seeks a $71 billion valuation

DeepSeek is in talks with investors for a new round at a pre-money valuation of roughly $71 billion — just a month after closing its first outside raise of about $7 billion — while beginning IPO preparations aimed at a mainland China listing. The Hangzhou lab says it needs the capital for gigawatt-scale data centers, in-house inference chips, and new AI agent products. A public debut would make China's highest-profile open-model lab one of the first frontier labs to test domestic capital markets.

Jul 8, 2026Top AI Stories

China's DeepSeek is quietly designing its own AI inference chip

Reuters reports that DeepSeek has spent about a year secretly recruiting chip engineers and courting manufacturing partners to build an in-house processor for AI inference — the stage where a trained model answers user queries. The goal is to cut its dependence on Nvidia and Huawei as US export controls keep tightening. It follows OpenAI's Broadcom-built Jalapeno inference chip and Anthropic's own reported chip ambitions — a sign that frontier labs increasingly want to own their silicon.

Jun 27, 2026Top AI Stories

DeepSeek open-sources DeepSpec, its speculative-decoding speedup stack

DeepSeek open-sourced DeepSpec, a full-stack codebase for training and evaluating speculative-decoding algorithms, alongside three drafting modules — DSpark, DFlash, and Eagle3 — that bolt onto its DeepSeek-V4 models to speed up text generation. Speculative decoding lets a small "draft" model propose several tokens at once for the larger model to verify in parallel; DeepSeek's paper reports generation speedups in the range of 60 to 85% on its V4 checkpoints. It is the latest in the lab's run of openly released inference tooling.

Jun 24, 2026Top AI Stories

Microsoft weighs hosting China's DeepSeek to cut Copilot Cowork costs

Microsoft told Axios it is exploring a self-hosted, fine-tuned version of China's DeepSeek V4 as a lower-cost engine for its new Copilot Cowork agent, possibly within weeks — an optional model run entirely on Azure, with added safeguards to reduce bias. The driver is surging inference costs, as agents make hundreds of model calls per task and OpenAI and Anthropic pull back from flat-rate pricing. In a Stratechery analysis, Ben Thompson casts the move as part of a broader pull toward China: Microsoft is strongly incentivized to use cheap, capable Chinese models, while memory makers Samsung, SK Hynix, and Micron may regret opening the door to Chinese chipmakers.

Jun 18, 2026Top AI Stories

US holds off blacklisting DeepSeek and 100-plus Chinese firms

Reuters reports that an interagency committee approved adding DeepSeek, memory-chip maker CXMT, and more than 100 other Chinese companies to the US Commerce Department's trade blacklist last year — but the Trump administration has held off to avoid escalating tensions with Beijing. A blacklisting would bar US firms from shipping them technology. Officials cite national-security concerns, including Anthropic's claim that DeepSeek tried to extract capabilities from Claude.

Jun 8, 2026Top AI Stories

DeepSeek nears a record $7.4 billion first funding round backed by Tencent and CATL

DeepSeek is close to sealing its first-ever outside funding round — about $7.4 billion, or 50 billion yuan — in one of China's largest startup financings. Tencent and battery maker CATL are the biggest external backers, alongside the state-backed National AI Industry Investment Fund and founder Liang Wenfeng, who is committing roughly $3 billion himself. The deal would value China's open-weights champion at $52 to $59 billion and signals Beijing's resolve to keep pace with the US capital surge.

May 7, 2026Top AI Stories

DeepSeek raising first VC round at $45 billion, more than double its valuation from weeks ago

DeepSeek is closing its first venture round at a reported **$45 billion valuation, up from $20 billion just weeks ago. The round is led by China Integrated Circuit Industry Investment Fund, with Tencent and Alibaba** participating. Founder Liang Wenfeng controls roughly 90% of the company and had not previously sought outside capital — the round is framed as a way to offer employee equity and retain talent against intensifying domestic competition. The valuation places DeepSeek alongside frontier US labs and reinforces China's push to build AI on Huawei silicon, independent of US export controls.

May 2, 2026Top AI Stories

DeepSeek ships V4 open-weights at 1.6 trillion params, 1 million-token context

DeepSeek released V4 last week and it's now landing on the local-LLM frontier. The flagship V4-Pro is a 1.6-trillion-parameter mixture-of-experts model with 49 billion active parameters and a 1-million-token context; smaller V4-Flash runs 284 billion total / 13 billion active. Both ship MIT-licensed on Hugging Face at $1.74 / $3.48 per million input/output tokens for Pro and $0.14 / $0.28 for Flash — undercutting GPT-5.4 Nano, Claude Haiku, and the Gemini variants.