Skip to main content
WireRead
Back to all news

Open-weight models

Kimi K3 and the open-weights escalation: what 2.8 trillion parameters actually buys

Moonshot published the largest set of open weights ever released. The benchmark story is genuinely strong; the deployment story is why the frontier labs can still sleep.

By , Editor-in-Chief · WireReadVerified September 2026

The answer

Moonshot published Kimi K3's 2.8-trillion-parameter weights on 27 July — the largest open release yet.

The open-weight story of 2026 has been one of steady encroachment rather than a single breakthrough: each Chinese release lands a little closer to the closed frontier, and each time the gap is measured in months rather than generations. Kimi K3 is the point at which the encroachment stopped being incremental. Moonshot AI launched it as a hosted service on 16 July and published the full weights on 27 July — 2.8 trillion total parameters, the largest open release ever made.

What was actually shipped

K3 is a Mixture-of-Experts model: 2.8 trillion parameters in total but only 104 billion active on any given token, with sixteen of 896 routed experts firing alongside two shared ones. That ratio is the design decision that makes the headline number liveable — you pay for 2.8 trillion parameters in storage and memory, but roughly 104 billion in compute per token. Moonshot pairs it with a 1,048,576-token context window and quantisation-aware training applied from the supervised fine-tuning stage onward, using MXFP4 weights and MXFP8 activations.

As Xinhua reported, a Moonshot AI executive explained the significance of the parameter count in simple terms: parameters are like neural connections in the human brain, and nearly 3 trillion of them means the model can "store more knowledge and patterns in its brain, understand more, think deeper, and answer more accurately."

Source: VentureBeat · 27 July 2026

That framing is Moonshot's, and it is worth treating as marketing rather than mechanism — parameter count has been a poor proxy for capability since at least 2023. The more informative detail is the quantisation-aware training, which means the model was taught to survive being compressed rather than compressed after the fact. For anyone who actually intends to serve it, that is the difference between a usable deployment and a degraded one.

Where it genuinely wins

The benchmark picture is better than open-weight sceptics expected and worse than the 'parity with the frontier' headlines suggested. Both halves are true at once, and the honest read requires holding them together.

Benchmark Kimi K3 Best closed comparison
Arena Frontend Code 1,679 (1st) Claude Fable 5, below K3
AA-Briefcase 1,527 (2nd) GPT-5.6 Sol Max 1,495
GDPval-AA v2 1,687 (3rd) Claude Fable 5 Max 1,815
FrontierSWE 81.2 Claude Fable 5 86.6
HLE-Full 43.5 Claude Fable 5 53.3

Winning Arena's Frontend Code evaluation outright, in blind developer testing, is not a soft result — it is humans preferring the open model's output on real work. Placing second on AA-Briefcase, a private agentic benchmark for long-horizon knowledge work, ahead of GPT-5.6 Sol Max, is similarly hard to dismiss. But Moonshot's own published comparison shows consistent losses to Claude Fable 5 across software engineering, reasoning and legal research, and losing on your own scorecard is about as candid as vendor benchmarking gets.

Why 'open' does not mean 'yours'

Two facts puncture the idea that anyone can now run frontier AI on their own terms. First, the hardware: Moonshot recommends at least 64 accelerators to serve K3, and a four-bit estimate puts the weights alone at roughly 1.4 terabytes. This is a datacentre artefact. Second, the licence — it ships under a custom Kimi K3 License, not MIT, despite widespread reporting to the contrary.

We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, deployment, and further innovation.

Source: Moonshot AI (Hugging Face model card) · 27 July 2026

The repository bears that out literally: 96 weight shards, a LICENSE file naming the Kimi K3 License, and the config needed to stand the model up. Moonshot has published the thing itself — on its own terms.

What open weights do buy is inspectability, the right to fine-tune, freedom from a single API's pricing and availability, and — for governments and regulated industries — the ability to run a frontier-class model inside their own perimeter. Those are substantial. They are not the same as democratisation, and the distinction matters when assessing what this does to the closed labs' pricing power.

The strategic read

Three things follow. The pricing floor keeps falling: K3 is listed at $3 per million cache-miss input tokens and $15 output, against $10/$50 for Claude Fable 5.1 — a model that beats it, but not by a margin every workload can justify paying five times for. Second, the weekend after K3, Alibaba announced a 2.4-trillion-parameter Qwen 3.8 with open weights, a model class it had historically kept API-only; the Chinese labs are escalating openness, not merely maintaining it. Third, and least comfortable: Anthropic accused Moonshot in February of training on 3.4 million Claude exchanges through distillation, and K3 benchmarks within a few points of the models named in that complaint. That allegation is unresolved, and it hangs over any straightforward reading of these results.

What to watch: whether Western enterprises actually deploy K3 in regulated environments, whether Moonshot's promised low and high reasoning-effort modes arrive (at launch it ran at maximum effort only, which made trivial prompts needlessly expensive), and whether the next closed release widens the verified-reasoning gap again or fails to.

Frequently asked questions

Is Kimi K3 the most powerful open-weight model?
It is the largest ever released at 2.8 trillion parameters and leads Arena's Frontend Code evaluation, but Moonshot's own tables show it behind Claude Fable 5 on several core benchmarks.
Can I run Kimi K3 myself?
Only with datacentre hardware. Moonshot recommends at least 64 accelerators, and the weights alone are roughly 1.4TB at four-bit precision.
What licence is Kimi K3 released under?
A custom Kimi K3 License, not MIT — a point widely misreported. Read the licence terms before commercial deployment.
How does Kimi K3's pricing compare to closed models?
It is listed at $3 per million cache-miss input tokens and $15 per million output, against $10/$50 for Claude Fable 5.1 — roughly a fifth of the cost on output.
What is the distillation dispute about?
Anthropic accused Moonshot in February 2026 of using 3.4 million Claude exchanges to train its models through distillation. The claim is unresolved and worth noting when reading K3's benchmark results.

Sources

← All news