← All guides
Kimi K3 Specifications: 2.8T Params, 1M Context & KDA

Kimi K3 Specifications: 2.8T Params, 1M Context & KDA

July 19, 2026 · kimi k3 specifications · kimi k3 specs · moonshot ai

Kimi K3’s spec sheet reads like a generational leap — because it is one. Here’s every major specification of Moonshot AI’s July 2026 flagship, explained in plain language. For the bigger picture, see our Kimi K3 complete guide.

Full specification table

SpecificationKimi K3
Total parameters2.8 trillion
ArchitectureKDA + Attention Residuals + Stable LatentMoE
Experts896 total, 16 active per token
Context window1,048,576 tokens (1M)
Vision / video inputNative
Long-context decode speedUp to 6.3× faster than Kimi K2
Scaling efficiency~2.5× better than Kimi K2
Weights licenseModified MIT — releases July 27, 2026
API model namekimi-k3

What the key specs actually mean

2.8 trillion parameters (MoE)

K3 is a mixture-of-experts (MoE) model. It doesn’t activate all 2.8T parameters for every token — instead, a router sends each token to just 16 of its 896 specialized experts. The result: the knowledge capacity of a massive model, with the inference cost of a much smaller one. This is why K3 can compete with GPT-5.6 Sol on benchmarks while costing about half as much per task.

1M-token context window

1,048,576 tokens is roughly 750,000+ words — entire codebases, full books, or long video transcripts in a single prompt. Context size alone isn’t new, but K3 pairs it with speed (below), which is what makes 1M tokens practical rather than a marketing number.

Kimi Delta Attention (KDA)

KDA is K3’s new attention mechanism, and it’s the reason long contexts are fast. Traditional attention gets slower quadratically as context grows. KDA changes how the model attends across long sequences, delivering up to 6.3× faster decoding at long contexts compared to K2 — so a 500K-token prompt doesn’t mean waiting minutes per response.

Attention Residuals

The second architectural breakthrough. Attention Residuals improve how information flows through the model’s layers, which Moonshot credits for K3’s ~2.5× better scaling efficiency — meaning K3 extracts far more capability per unit of training compute than K2 did.

Stable LatentMoE

The routing system that keeps 896 experts training and serving stably. More experts means more specialization (better quality), but historically harder training stability; Stable LatentMoE is Moonshot’s solution, allowing the jump from K2’s 384 experts to 896.

Native vision and video input

Unlike K2, K3 accepts images and video natively — no separate adapter model. This shows up in K3’s visual-agent benchmark results, like its #1 BrowseComp score (91.2).

How K3’s specs compare to K2

SpecKimi K2Kimi K3
Total parameters1.0T2.8T
Experts384896 (16 active)
Context window256K1M
Vision / videoNative
Frontend Code Arena#18 (K2.6)#1

We break down which model wins for which use case in Kimi K3 vs Kimi K2: Full Comparison.

Can you self-host these specs?

Honestly — not easily. 2.8 trillion parameters, even in MoE form with 16 active experts, requires serious multi-GPU infrastructure. When the weights drop on July 27 under the Modified MIT license, self-hosting will be realistic for well-resourced labs and companies, not hobbyists. Everyone else should use the Kimi app, Kimi Work desktop app, or the API at platform.moonshot.ai.

FAQ

How many parameters does Kimi K3 have? 2.8 trillion total, in a mixture-of-experts design with 16 of 896 experts active per token — roughly 50B active parameters per token.

What is Kimi Delta Attention (KDA)? K3’s new attention mechanism — a linear-attention variant enabling up to 6.3× faster long-context decoding than K2, with a prefill-cache implementation contributed to vLLM.

What is Kimi K3’s context window? 1,048,576 tokens (1M) — entire codebases, books, or long video transcripts in one prompt.

Can I self-host Kimi K3? The weights release July 27, 2026 (Modified MIT, MXFP4), but Moonshot recommends supernode setups with 64+ accelerators. Most users should use the app or API.

The takeaway

Kimi K3’s specifications aren’t just bigger numbers — the KDA + Attention Residuals architecture makes its headline features (1M context, 896 experts) usable in the real world. That’s why it ranks #4 of 189 models on Artificial Analysis, the highest ever for an open-weight model. Want to see what those specs deliver in practice? Start with the complete Kimi K3 guide or try it free in the Kimi app.

Try it yourself

Sign up for Kimi through our invite link and both of us get free bonus membership credits — up to a full year, at no cost to you.

Claim Free Credits →