Every published Top AI Stories item tagged with Alibaba Cloud, newest first.
A new Anthropic threat intelligence report attributes distillation campaigns with high confidence to Alibaba, Moonshot AI, DeepSeek, Zhipu, Xiaomi, SenseTime and MiniMax. A joint advisory from three US agencies asserted this pattern two days earlier; what is new here is the platform's own telemetry. Alibaba ran more than 151 million exchanges between May and July, peaking near three million a day across thousands of fraudulent accounts, to capture the chain-of-thought reasoning behind Opus. Moonshot and DeepSeek went further, silently routing their own customers' requests to Claude and keeping the transcripts, exposing names, live credentials and corporate data to a third party those users never chose.
The National Security Agency, the Federal Bureau of Investigation and the Cybersecurity and Infrastructure Security Agency issued a joint advisory saying DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.ai have extracted billions of tokens from Claude, GPT, Gemini and Grok since late 2024, likely with Chinese government awareness. That escalates a July accusation from the White House technology chief, which named Moonshot alone. Distillation, training a smaller model on a larger one's outputs, is a standard technique — but the agencies say it forms the core rather than a supplement of these labs' development. They also reject the number behind the 2025 market shock, DeepSeek's widely cited training cost of 5.6 million dollars. The advisory says it leaves out the cost of the data taken this way. China's foreign ministry calls the claims groundless.
Two days after saying the weights would land that evening, Alibaba's Qwen team put Qwen3.8-Flash-Next on Hugging Face — ungated, 132 model files, a mixture-of-experts (MoE) design with 125 billion total parameters that activates roughly 6 billion per token. It is the second Alibaba release this month to pair open weights with a carve-out for service providers, and this one is drawn tighter. The Qwen Community License 1.0 permits commercial use, hosting and fine-tuning, but anyone running a model-as-a-service business, or an AI coding or office assistant, must get a separate agreement from Alibaba at any size. There is no revenue floor on that clause. The floor applies only to a different requirement, to display the model's name on screen, which starts at 100 million monthly active users or $20 million in monthly revenue.
Alibaba's Qwen team said it will release Qwen3.8-Flash-Next, a multimodal mixture-of-experts (MoE) model with 125 billion total parameters that activates roughly 6 billion of them per token, at 11 in the evening Beijing time on Wednesday. The framing is the unusual part. This is a technology preview of the architecture that will carry the full Qwen 4 family rather than a flagship in its own right, so developers get to build against the new design months ahead of the models that will use it. At the time of writing the weights were not yet on Hugging Face, no benchmark scores had been published, and no license had been named.
Nari Labs published how it optimized Alibaba's Qwen3-TTS CustomVoice, a 1.7 billion parameter text-to-speech model, to return its first audio in under 50 milliseconds at the 95th percentile and stay under 100 milliseconds at twenty requests a second. The techniques are unglamorous — one scheduler across the model's three modules, trimming leading silence, cached incremental decoding — and the implementation is on GitHub. At roughly two dollars per million characters on a single H100, real-time voice moves within reach of small teams.
Qwen3.8-27B landed on Hugging Face with open weights, no gating, and native image and video input. Its model card reports 61.7 on SWE-bench Pro and 84.3 on the OSWorld computer-use benchmark, with a native context window of 262,144 tokens that extends to one million. The license matters as much as the scores: the flagship Qwen3.8-Max shipped days earlier under custom terms carrying a $50 million revenue trigger, while this one is unrestricted Apache 2.0.
Alibaba put the weights for Qwen3.8-Max on Hugging Face — a mixture-of-experts model with 2.4 trillion total parameters and about 95 billion active per request, natively handling 262,144 tokens of context. The license is not open source. Any company running a model-as-a-service or AI work assistant business whose revenue tops $50 million over twelve consecutive months must obtain a separate license from Qwen before using the model commercially, and products above 100 million monthly active users must display the model name prominently. Reuters had reported a revenue share; the published terms instead require a negotiated agreement.
The Bureau of Industry and Security, the US Commerce Department's export-enforcement arm, is reviewing how Chinese AI companies reach Nvidia hardware by renting computing power housed in other countries — an arrangement that is not currently illegal under the export rules. Southeast Asia is the focus, with operators in Thailand, Malaysia and Singapore leasing chips to Chinese clients; the Singaporean firm Megaspeed is already under investigation. Nvidia, which has said the controls already cost it the world's second-largest market, would fight any limit on overseas data-center access.
Reuters reports that Alibaba plans to require large commercial users of the open-weight Qwen3.8-Max to hand back a share of the revenue they earn from it, arriving alongside a weights release expected within days. The model carries 2.4 trillion total parameters and 95 billion active ones, and the rate is still being negotiated. It follows Moonshot AI's Kimi K3 license, which makes anyone selling the model as a service above $20 million in annual revenue sign a commercial agreement — more evidence that Chinese open weights now arrive with commercial strings attached.
Qwen3.8-Max is a mixture-of-experts model with 2.4 trillion total parameters, roughly 95 billion of them active per request, and a context window of about one million tokens. Alibaba reports 86.6 on Terminal-Bench 2.1 — ahead of Claude Opus 4.8 and Claude Fable 5 at 84.6, behind GPT-5.6 Sol at 88.8 — while trailing Fable 5 badly on SWE-bench Pro, at 67.7 against 80.0. The hosted API is live now at $2 per million input tokens and $6 per million output tokens. Open weights for Max and the smaller Qwen3.8-27B are promised for next week and are not published yet.
Nearly 200 venture-backed startups, organized by the newly formed Little Tech Association, sent a letter urging the Trump administration not to ban US access to Chinese open-weight AI models. Signatories including Y Combinator and Proton argue that cheap open models from Moonshot AI and Alibaba are a lifeline for small companies, and that a broad prohibition would raise costs and hand even more power to a few dominant US labs. They asked for "targeted safeguards" rather than a blanket cutoff.
In this week's Stratechery, Ben Thompson argues that Western alarm over cheap Chinese open-weight models — Moonshot's Kimi K3 and Alibaba's forthcoming Qwen3.8 Max — misreads the economics. What matters, he writes, is not the sticker price per token but the total cost of a correct answer, since different models burn very different token volumes to solve the same problem. As intelligence becomes a commodity, he contends, US frontier labs with better cost structures still win — and he urges loosening rules that push American security teams toward Chinese models.
Alibaba's Qwen team unveiled Qwen-Image-3.0, a text-to-image model built to render dense, information-heavy pictures — newspaper pages, multi-panel infographics, even math-filled academic papers — from prompts up to 4,500 tokens, with native text in twelve languages. Notably, the release carried no benchmarks, no parameter count, no license, and no downloadable weights, a sharp break from the open, Apache-licensed launches of Qwen-Image 1.0 and 2.0. It is an early sign that even China's most open lab may be closing up its flagship image model.
Alibaba's Qwen team unveiled a preview of Qwen 3.8 Max, its first multimodal model above one trillion parameters, claiming it trails only Anthropic's Claude Fable 5 among frontier systems. The 2.4-trillion-parameter mixture-of-experts model handles text, images, video, and documents, and Alibaba says it beats its predecessor on coding, full-stack development, and office workflows. But the company published no benchmark table, no open weights, and — critically for a sparse model — never disclosed how many parameters are active per token, the figure that determines real serving cost. The preview arrived two days after Moonshot's 2.8-trillion-parameter Kimi K3, underscoring how fast Chinese labs are racing at the open-weight frontier.
China's Cyberspace Administration has cleared Apple Intelligence for release, ending a roughly two-year wait since the features debuted in the United States. Alibaba's Qwen model will handle text and image understanding and generation across iOS, iPadOS, macOS, and visionOS for Chinese users, with Baidu contributing a smaller model alongside it. Apple had explored deals with Baidu, DeepSeek, and ByteDance before settling on Alibaba. Approval is not the same as availability — a launch is expected to track Apple's usual autumn software cycle.
China's Ministry of Commerce has held meetings with Alibaba, ByteDance, and startup Zhipu AI about whether to limit foreign access to the country's most capable models, according to Reuters — a striking reversal for labs whose open-weight releases have been the main challenge to US frontier dominance. Options sketched in the talks reportedly range from security reviews to barring the most sensitive models from public release, and cover open-weight systems like Qwen, Doubao, and GLM, not just proprietary ones. Nothing is decided, and officials have made no public comment.
Alibaba will bar employees from using Anthropic's Claude Code starting July 10, after Chinese outlet Yicai reported the coding assistant carried what it called embedded "backdoor" risks. Developers had flagged that Claude Code inspected user environments — checking timezone and proxy details and inserting subtle markers into prompts sent to Anthropic; Anthropic says that was a March anti-abuse experiment to stop unauthorized resellers and model distillation, not surveillance. Staff are being pointed to Alibaba's own Qoder tool instead, deepening a months-long feud between the two labs.
Anthropic told the US Senate Banking Committee, in a June 10 letter, that operators tied to Alibaba and its Qwen AI lab ran roughly 25,000 fraudulent accounts to query Claude more than 28.8 million times between April 22 and June 5 — what Anthropic calls the largest known "distillation" attack against it, a technique that trains a cheaper rival model on a stronger one's outputs. It echoes earlier campaigns Anthropic attributed to DeepSeek, Moonshot, and MiniMax, and sharpens the US-China fight over who gets to copy frontier AI.
Researchers on Alibaba's Qwen team released Qwen-AgentWorld, a pair of open models — 35 billion and 397 billion parameters — that simulate entire software and tool environments in language, predicting how a world changes in response to an agent's actions. The team uses them two ways: as cheap simulators for training agents with reinforcement learning, and as a pre-training foundation that lifts agent scores across seven benchmarks. Strikingly, agents trained inside the simulated world beat those trained in the real environment alone, hinting that model-generated environments could become central to building capable agents.
Alibaba's Tongyi Lab released Qwen-Robot, a suite of three foundation models that give robots the software to navigate spaces, manipulate objects, and predict how the physical world will respond. The manipulation model was trained on more than 38,000 hours of data and topped a leading robotics benchmark. The launch pushes Alibaba's open-model strategy from chatbots into embodied AI, where Chinese firms are racing to build a common operating layer for the coming wave of humanoid and warehouse robots.
Reuters, citing Bloomberg News, reports that Beijing has widened informal travel restrictions originally placed on senior DeepSeek researchers to AI talent at Alibaba and other private firms — requiring some professionals to seek official approval before traveling abroad, with the policy framed around state-secret concerns and strategically important AI work. The move treats AI researchers themselves as restricted assets, a national-security frame that mirrors the logic the US has used in reverse for chip export controls. The pattern hardens a two-way decoupling at the human-capital layer rather than just the supply chain.