Moonshot AI’s Kimi K3 Is Now the Largest Open-Weight AI Model Ever Released

Moonshot AI published full open weights for Kimi K3 on July 27, 2026, and the number alone is the story: 2.8 trillion parameters, making it the largest open-weight AI model ever released — bigger than anything DeepSeek, Alibaba’s Qwen team, or Meta’s Llama line has shipped publicly. But the more interesting number is 104 billion. That’s how many of those 2.8 trillion parameters actually fire on any given token, and it’s why this is as much an efficiency story as a scale story.

What Moonshot Actually Shipped

Kimi K3 is a mixture-of-experts model: 896 individual “experts” sit inside the architecture, but only 16 activate per token, which is how Moonshot gets a 2.8T-parameter system to run at roughly the cost of a much smaller dense model. The company pairs this with two custom techniques — “Kimi Delta Attention” and “Attention Residuals” — that it claims deliver about 2.5x better scaling efficiency than a standard dense architecture. The model also ships with native multimodal vision support and a 1-million-token context window, putting it in the same league as frontier closed models on raw capability specs.

One nuance worth getting right: Kimi K3 is open-weight, not open-source in the strict sense. Moonshot released it under a custom “Kimi K3 License” that carries commercial-scale conditions — meaning anyone can download and run the weights, but there are strings attached once you’re operating at real scale. Calling it “fully open” would overstate what Moonshot actually did.

Data center server racks behind a glass partition, representing the large-scale compute infrastructure used to train and run models like Kimi K3
Photo: Cory M. Grenier, CC BY-SA 2.0

How It Stacks Up Against Everything Else

Context matters here. Every previous “largest open-weight model” — DeepSeek V3 and V4, Alibaba’s Qwen line, Meta’s Llama 4 Maverick — topped out in the hundreds of billions of parameters. Kimi K3’s 2.8 trillion is a genuine step change, reportedly the first open-weight release to approach the 3-trillion mark. Alibaba’s newer Qwen3.8-Max-Preview is currently hosted-only, with an open release still pending, so Moonshot’s “largest ever” claim holds for now — though it may not hold for long.

On raw performance, early reporting puts Kimi K3’s strongest results in coding and math, where it’s competitive with closed models in the GPT-4.1 and Claude 3.7 Sonnet performance class. Moonshot itself is reportedly upfront that K3 still trails the very top closed frontier systems overall. That’s a more honest framing than the headline number suggests, and it’s worth holding onto if you see bolder benchmark claims circulating elsewhere — several secondary write-ups cite comparisons that couldn’t be independently verified, so treat anything beyond “competitive in coding and math, behind the true frontier” with some skepticism.

Why This Matters Beyond the Spec Sheet

The practical hook for readers isn’t parameter count — it’s what open weights let you do. Because Kimi K3 can be self-hosted, users and companies wary of routing data through a Chinese-hosted API get an alternative: run the model on your own infrastructure and keep your data off Moonshot’s servers entirely. That data-sovereignty angle is quietly becoming one of the bigger reasons open-weight releases matter in 2026, separate from how the model scores on a leaderboard.

It also fits a pattern that’s been building all year: Chinese labs — DeepSeek, Qwen, Zhipu’s GLM, and now Moonshot’s Kimi line — have driven most of the genuinely open frontier releases, while the leading US labs mostly keep their best models closed. For contrast, see how Anthropic approached its own flagship release with Claude Opus 5 — a closed system built for enterprise scale, not for anyone to download and self-host. That divergence is exactly why several Western policy circles are reportedly renewing calls for regulation around the wave of Chinese open-weight releases — a fast-moving, unrestricted model is a very different policy problem than a closed one behind an API.

What’s Next

Don’t expect Kimi K3 to hold the “largest open-weight model” title uncontested for long — Alibaba’s Qwen team has already signaled an open release is coming, and the pace of this race has been closing every few months, not years. The bigger story is the compute and infrastructure race sitting underneath all of it: training and serving models at the trillion-parameter scale is exactly what’s driving the current spending spree across the AI industry. We’ve covered the investor doubts swirling around Big Tech’s AI infrastructure bet in detail, and Kimi K3’s release is one more data point in why that spending keeps climbing.

For now, the weights are live on Hugging Face, and anyone with the hardware to run a trillion-parameter MoE model can try it today — that’s the part of this story that closed frontier labs still can’t match.

Deixe um comentário

O seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *