← All guides
Kimi K2.8 Preview: 1M Context, Plans, Pricing & K2.8 vs K3 vs K2.7

Kimi K2.8 Preview: 1M Context, Plans, Pricing & K2.8 vs K3 vs K2.7

September 25, 2026 · kimi k2.8 preview · kimi k2.8 · kimi code · kimi-for-coding · kimi k2.8 vs k3 · kimi k2.7 code · kimi code models · kimi code pricing

Kimi K2.8 Preview is the model that now answers when you call kimi-for-coding in Kimi Code. Moonshot swapped it in on September 11, 2026 in place of K2.7 Code, with no ID change: every plan that includes Kimi Code now gets a 1M-token context window (1,048,576 tokens), image and video input, and K3-style low / high / max thinking. Moonshot says performance is “close to K3”. It has not published parameters, benchmarks, open weights or a per-token API price.

This guide sorts what is confirmed from what is marketing or rumor, as of September 25, 2026. It covers the exact tier rules (including the new Go/Plus/Pro/Max plans launched on September 18), whether you can reach K2.8 through an API, the “thinking off” routing trap, and the question most people actually have: K2.8 Preview, K3 or K2.7 Code HighSpeed, and which one for which job?

Our primary sources are Moonshot’s own Kimi Code changelog, the Kimi Code model configuration page, the membership page and the Open Platform price table, checked against news coverage and third-party listings.

What Kimi K2.8 Preview actually is

The official announcement is a single changelog entry dated September 11. It makes three claims: coding and agent capability “improved across the board” with “significantly more efficient thinking than K2.7 Code”; the same thinking levels as K3 (default max); and “a context window of up to 1M is available across all membership tiers.” Pandaily and TechFlow via KuCoin reported the same facts the same day. Pandaily adds that secondary reports put the model in Kimi Work too. We could not confirm that from a Moonshot page, and we found no standalone launch thread from @Kimi_Moonshot.

Here is the confirmed-versus-unknown ledger:

QuestionStatus (Sep 25, 2026)Source
Release dateSeptember 11, 2026 (full rollout in Kimi Code)Kimi Code changelog
Model IDkimi-for-coding (unchanged; the name “K2.8 Preview” fails as an ID)Model configuration page
Context window1,048,576 tokens on all Kimi Code plansChangelog + model table
Input typesText, image and videoModel table
Thinkinglow / high / max, default maxModel table
Speed classRegular (not HighSpeed)Model table
Parameters / architectureNot disclosedn/a
BenchmarksNone published (“close to K3” only)n/a
Open weightsNone; no K2.8 repo under moonshotai on Hugging FaceHF model list
Open Platform (pay-as-you-go) APINot listedPlatform model list
Stable (non-preview) dateNot announcedn/a

Two claims circulating in news write-ups are wrong, so skip them. K2.8 Preview does not “introduce vision to the Kimi family”: the K2.7 Code model card already lists a 400M-parameter MoonViT vision encoder, and K3 is natively multimodal. And it is not an Open Platform API model. Right now it lives inside Kimi Code.

The Kimi Code lineup after the swap

Kimi Code now offers three models under four IDs. This is the table that matters for anyone using the CLI, VS Code extension, the new Desktop app (shipped September 17) or a third-party agent with a Kimi Code key:

Model IDModelContextThinking (default)InputSpeed / quota noteNew plansLegacy plans
k3Kimi K3Up to 1,048,576low / high / max (high)Image, video~2× the quota of k3-256kPlus+ (1M on Pro+)Moderato+ (1M on Allegretto+)
k3-256kKimi K3262,144 fixedlow / high / max (high)Image onlyRegularPlus+Moderato+
kimi-for-codingK2.8 Preview1,048,576low / high / max (max)Image, videoRegularPlus+Andante+
kimi-for-coding-highspeedK2.7 Code HighSpeed262,144Always onImage, video~6× speed, 3× quotaPro+Allegretto+

Source: Kimi Code model configuration, fetched September 25, 2026.

Two things jump out. First, standard K2.7 Code is gone from Kimi Code. Only its HighSpeed variant survives, and Moonshot describes that one as “the same coding ability” as K2.7 Code at ~5–6× the output speed. Second, the defaults differ: K3 defaults to high effort, K2.8 Preview to max. “More efficient thinking” at max effort is not automatically cheaper per turn than K3 at high effort, and Moonshot publishes no quota multiplier for K2.8 relative to K3.

Who gets 1M context now: the tier map

“All membership tiers” needs one footnote. It means every tier that includes Kimi Code, not literally every Kimi subscription. On September 18 Moonshot relaunched memberships under new names after pausing new subscriptions in July for capacity reasons. According to KuCoin’s report, annual prices are unchanged:

New planMonthly (annual)Kimi Code?K2.8 Preview (1M)K3K3 1MHighSpeed
Go49 CNY (468/yr)No coding quotaNoNoNoNo
Plus99 CNY (948/yr)YesYesYes (256K max)NoNo
Pro199 CNY (1,908/yr)YesYesYesYesYes
Max699 CNY (6,708/yr)YesYesYesYesYes

Monthly-billed prices from Kimi’s membership pricing page; annual prices as reported by KuCoin/TechFlow (annual works out to 39/79/159/559 CNY a month). Your regional price may differ. Entitlements from the membership page and model table.

Legacy subscribers keep their old plans, and the docs pair them with the new ones: Moderato with Plus, Allegretto with Pro. The one asymmetry is at the bottom. Legacy Andante members still get Kimi Code, and K2.8 Preview is the only model they can call, while new Go members get no coding quota at all.

The practical consequence is the most interesting fact in this release. On Plus or Moderato, K2.8 Preview has a bigger context window than K3. K3 is capped at 256K on those plans, so for a mid-tier member who needs to load a large repository, a long log dump or a video into one session, K2.8 Preview is the only 1M option. Before September 11 the answer was “upgrade to Allegretto.”

Quota rules also changed for new members. The weekly limit is gone, but the rolling 5-hour rate window and a monthly total quota (shared with Kimi on the web) remain. Legacy members keep the 7-day refresh with no rollover. When quota runs out, the optional Extra Usage balance takes over at rates Moonshot says are “close to the official API pricing of the Kimi Open Platform,” billed in RMB with a ¥25 minimum top-up.

Is K2.8 Preview on the API? What it costs

This is where the confusion lives, because “API” means two different things at Moonshot.

  1. The Kimi Code API is a membership-keyed endpoint for coding tools. Base URLs are https://api.kimi.ai/coding/v1 (OpenAI protocol) and https://api.kimi.ai/coding/ (Anthropic protocol) overseas, with api.kimi.com equivalents in China. K2.8 Preview is available here as kimi-for-coding, billed against your plan quota rather than per token. It is designed for agents such as Kimi Code CLI, Claude Code, OpenCode and Codex, not for building a production backend.
  2. The Open Platform API (api.moonshot.ai/v1, now documented at platform.kimi.ai) is pay-as-you-go. K2.8 is not on it. The price table lists only the four models below:
Open Platform modelCached inputInputOutputContext
kimi-k3$0.30$3.00$15.001,048,576
kimi-k2.7-code$0.19$0.95$4.00262,144
kimi-k2.7-code-highspeed$0.38$1.90$8.00262,144
kimi-k2.6$0.16$0.95$4.00262,144
K2.8 PreviewNot listedNot listedNot listedn/a

Per 1M tokens, from the official price table on September 25, 2026. K3 additionally bills cache writes at $3 (5-minute TTL) or $6 (1-hour TTL) per 1M.

OpenRouter’s Moonshot page has no K2.8 model either. Two resellers do list one: ZenMux shows moonshotai/kimi-k2.8-preview at $1 input / $4 output / $0.25 cache read, and APIMaster advertises the same $1/$4. APIMaster itself notes that Moonshot “do[es] not publish a separate K2.8 pay-as-you-go token table.” Neither reseller explains how it sources a model that Moonshot’s public API does not sell, so treat those endpoints as unofficial. Don’t put anything you can’t afford to lose behind them. For scale, $1/$4 sits right next to K2.7 Code’s official $0.95/$4 and roughly a third to a quarter of K3’s $3/$15.

ZenMux also shows its own measured 41.1 tok/s throughput and 1.43 s latency for the model. That is one reseller’s live metric, not a benchmark, but it’s the only speed figure we found. For context, K3’s standard speed on Moonshot’s API is about 34 tok/s. If you’ve seen “260 tokens/s” attached to K2.8, that figure belongs to K2.7 Code HighSpeed: the platform model list says it runs at “about 180 tokens/s, up to 260 tokens/s in short-context scenarios.”

The “thinking off” routing trap

One line in the docs deserves more attention than the context headline:

“With thinking turned off, requests to the K3 series and K2.8 Preview are served by K2.8 Preview (no thinking).”

If your tool sends none as the effort, or disables thinking, while you have k3 selected, you are not talking to K3. You get K2.8 Preview without reasoning. That matters for evals, cost attribution and bug reports. If a “K3” session suddenly feels different, check the effort setting first.

Kimi Code maps third-party effort strings as follows:

Tool sendsKimi Code applies
ultra, max, xhighmax
high, mediumhigh (recommended)
low, minimum, lightlow
nonethinking disabled (routes to K2.8 Preview, no thinking)
null / undefinedmodel default (high for K3, max for K2.8 Preview)
anything elseHTTP 400

Three more gotchas from the same page:

  • Switching models or effort invalidates the context cache. Usage spikes right after a switch because the prefix has to be re-prefilled. Moonshot’s advice is to pick one effort per session and start a new session when you change models.
  • A mistyped HighSpeed ID fails silently. Anything other than exactly kimi-for-coding-highspeed falls back to standard kimi-for-coding, which is now K2.8 Preview. You get no error and no speedup.
  • k3-256k does not accept video. If your session history contains video, compact before switching down from k3 or the switch fails.

K2.8 Preview vs K3 vs K2.7 Code: what’s actually different

Putting every verified number side by side:

K2.8 PreviewKimi K3K2.7 Code (HighSpeed in Kimi Code)
ReleasedSep 11, 2026Jul 16, 2026Jun 12, 2026
SizeUndisclosed2.8T total, 104B active MoE1T total, 32B active MoE
Max context1M on all Kimi Code plans1M (Pro/Allegretto+ in Kimi Code)256K
Thinkinglow / high / max, default maxlow / high / max, default highAlways on
InputImage, videoImage, video (k3-256k: image only)Image, video
Open weightsNoYes (details)Yes (Modified MIT)
Public benchmarksNoneYes (vendor + third-party)Yes (e.g. Kimi Code Bench v2 62.0, MCP-Atlas 76.0)
Pay-as-you-go APINo (Kimi Code only)$0.30 / $3 / $15$0.19 / $0.95 / $4 (HighSpeed $0.38 / $1.90 / $8)
Kimi Code quota costNot published1M ≈ 2× the 256K variantHighSpeed ≈ 3× quota for ~6× speed
Moonshot’s positioning”Close to K3”, good at completion and routine dev”Most capable flagship coding model”Same ability as K2.7 Code, 5–6× faster output

K2.7 Code figures come from its Hugging Face model card and the changelog, which also cites +10.4% on Program-Bench and 30% fewer reasoning tokens than K2.6. K3 figures follow our specifications page.

Read the table carefully and there is no K2.8 number on the quality axis at all. “Close to K3” is Moonshot’s own assessment of an unreleased-size preview model. It may well be true. But until Moonshot or an independent evaluator publishes scores, the only evidence you can trust is your own repository.

Which should you pick?

The right answer depends mostly on your plan, then on the job.

If you’re on Plus / Moderato (or legacy Andante):

  • Default to K2.8 Preview (kimi-for-coding). It is the only 1M-context option you have, it takes video, and Moonshot positions it for completion and routine development.
  • Use k3-256k for the hardest reasoning-heavy tasks that fit in 256K: a gnarly bug across a few modules, an architecture decision, a tricky kernel. (Andante has no K3 access at all.)

If you’re on Pro / Allegretto or above:

JobPickWhy
Everyday fixes, tests, refactors, PR reviewsK2.8 Preview”Routine development” is its stated sweet spot; efficient thinking
Tight edit-run loops where output speed dominatesK2.7 Code HighSpeed~6× output speed, but 3× quota and 256K cap
Long-horizon agentic work, multi-hour goals, whole-repo migrationsK3 (k3)Flagship, 1M, strongest published results
Very long context, but cheap on quotaK2.8 Preview1M without the ~2× quota of k3 at 1M
Hard problem that fits in 256Kk3-256kFlagship quality at roughly half the 1M quota
Anything with video input and >256K historyk3 or K2.8 Previewk3-256k rejects video

If you need a production API rather than a coding agent: K2.8 Preview is not for you yet. Use kimi-k3 for quality, or kimi-k2.7-code at $0.95/$4 for cost, on the Open Platform or one of the third-party hosts. Both also have open weights, which K2.8 does not.

If you pin kimi-for-coding in CI or scripts: your pipeline changed models on September 11 without a config change. Re-run your evals, watch token usage (the default effort is now max), and consider pinning an explicit effort level. Pandaily flags the same risk, and Moonshot hasn’t said whether the ID will move again when “Preview” becomes stable.

Honest caveats

  • “Preview” is doing real work in the name. There’s no GA date, and Moonshot could change behavior, quota cost or the model behind kimi-for-coding again without warning, as it just did.
  • No numbers means no ranking. We don’t know where K2.8 sits against K2.x-to-K3 gains on any benchmark. Anyone quoting K2.8 scores as of this writing is quoting something Moonshot hasn’t published.
  • The consumer app is unconfirmed. Moonshot’s docs only describe Kimi Code (CLI, VS Code, Desktop, third-party tools). Pandaily’s Kimi Work mention cites secondary reports.
  • Open weights are uncertain. Moonshot open-sourced both K2.7 Code and K3, which is a reason to expect a K2.8 release eventually, but a preview that’s closed today is not a promise.
  • Reseller endpoints are unofficial. The $1/$4 listings are the only per-token prices in circulation, and Moonshot hasn’t documented them.

How to switch to K2.8 Preview

If you already use Kimi Code with defaults, you’re on it. Otherwise:

  • Kimi Code CLI: /model and pick kimi-for-coding. If it isn’t listed, /logout then /login.
  • VS Code extension: choose it from the model dropdown in the input bar (restart VS Code if missing).
  • Kimi Code Desktop: model picker inside the Composer.
  • Claude Code, OpenCode, Codex: create a key in the Kimi Code Console, set the base URL above and model ID kimi-for-coding. Our Kimi Code setup guide walks through the config.

Start a fresh session after switching so you aren’t paying to re-prefill a cache that won’t hit.

FAQ

What is Kimi K2.8 Preview? Kimi K2.8 Preview is Moonshot AI’s new mid-tier coding model, rolled out in Kimi Code on September 11, 2026. It replaced K2.7 Code behind the unchanged model ID kimi-for-coding, supports low/high/max thinking (default max), accepts image and video input, and has a 1M-token (1,048,576) context window on every plan that includes Kimi Code. Moonshot describes its performance as close to K3 with more efficient thinking than K2.7 Code, but has not published parameters, benchmarks or weights.

Is Kimi K2.8 Preview available on the API? Only through Kimi Code. As of September 25, 2026, Moonshot’s Open Platform model list and price table show kimi-k3, kimi-k2.7-code, kimi-k2.7-code-highspeed and kimi-k2.6, but no K2.8 model. You can call K2.8 Preview as kimi-for-coding through the Kimi Code API (api.kimi.ai/coding/v1 or the Anthropic-compatible api.kimi.ai/coding/) with a membership API key. A couple of third-party resellers list it at $1/$4 per million tokens, but Moonshot has not documented that route.

Which Kimi plans get K2.8 Preview with 1M context? Every plan that includes Kimi Code: Plus, Pro and Max on the new plans, and Andante and above on legacy plans. The new entry-level Go tier has no coding quota at all. That makes K2.8 Preview the only way a Plus or Moderato member can get a 1M window in Kimi Code, because K3 is capped at 256K on those plans and needs Pro or Allegretto for 1M.

Is Kimi K2.8 Preview better than Kimi K3? No. Moonshot’s own wording is performance close to K3, and its model table still calls K3 the most capable flagship coding model. K2.8 Preview’s advantages are availability and efficiency: 1M context on lower plans and more efficient thinking. For hard, long-horizon engineering work, K3 remains the stronger pick when your plan allows it.

How much does Kimi K2.8 Preview cost? Inside Kimi Code it is included in the membership quota, with no separate per-token price. The new plans start at 99 CNY a month for Plus billed monthly (948 CNY a year, about 79 a month), the cheapest tier with Kimi Code. Moonshot has not published a pay-as-you-go K2.8 rate. For comparison, K2.7 Code on the Open Platform costs $0.19 cached input, $0.95 input and $4 output per million tokens, and K3 costs $0.30, $3 and $15.

Is Kimi K2.8 Preview open source? Not so far. Moonshot’s Hugging Face organization shows Kimi-K3 and Kimi-K2.7-Code repositories but no K2.8 repository as of September 25, 2026, and no license or model card has been published. Both K3 and K2.7 Code shipped with open weights, so a later release is plausible, but nothing has been announced.

What happened to Kimi K2.7 Code? In Kimi Code, standard K2.7 Code was replaced in place: the kimi-for-coding ID now serves K2.8 Preview. K2.7 Code HighSpeed stays available as kimi-for-coding-highspeed on Pro and Allegretto plans and above. Outside Kimi Code, kimi-k2.7-code is still listed on Moonshot’s Open Platform and its open weights remain on Hugging Face.

How do I use K2.8 Preview in Claude Code or other third-party tools? Create an API key in the Kimi Code Console, point the tool at the Anthropic-compatible base URL (https://api.kimi.ai/coding/ overseas) or the OpenAI-compatible one (https://api.kimi.ai/coding/v1), and set the model ID to kimi-for-coding. Use the ID, not the name K2.8 Preview, which fails. Start a new session after switching, because the old context cache will not hit on the new model.

The bottom line

K2.8 Preview is a plan-tier story more than a model story. Moonshot hasn’t shown a single benchmark, but it has made 1M context the default for every Kimi Code subscriber. On Plus and Moderato that means the cheaper model now out-reaches the flagship on context length. Use it as your daily driver for routine coding. Keep K3 for the long-horizon, high-stakes work it was built for, and reach for K2.7 Code HighSpeed only when output speed is the bottleneck and you can spare 3× quota. If you need a real per-token API or open weights, K2.8 isn’t there yet: stay on kimi-k3 or kimi-k2.7-code.

Keep reading: the complete Kimi K3 guide, Kimi K3 specifications, every Kimi K3 API provider, Kimi K3 vs Kimi K2, and how to get Kimi on your PC.