← All guides
Kimi K3 vs DeepSeek V4.1 Flash: Price, Benchmarks & When to Switch

Kimi K3 vs DeepSeek V4.1 Flash: Price, Benchmarks & When to Switch

September 25, 2026 · kimi k3 vs deepseek · deepseek v4.1 flash · deepseek flash pricing · deepseek v4 flash 0731 · kimi k3 pricing · kimi k3 benchmarks · open weight llm comparison

DeepSeek V4.1 Flash is 10 to 25 times cheaper per token than Kimi K3, about six times faster, and beats K3 on most of the agentic coding benchmarks DeepSeek published. Kimi K3 still leads on independent general-intelligence scoring (44 vs 39 on Artificial Analysis), on reasoning tests like GPQA Diamond and Humanity’s Last Exam, and on maximum output length. For cost-sensitive agent workloads, Flash is now the model to test first; K3 is the one to keep for your hardest tasks.

This comparison is current as of September 25, 2026, fifteen days after DeepSeek released V4.1 Flash and retired V4-Flash-0731 on its API. It covers specs, what changed since 0731, advertised versus real cost, licenses, verified benchmarks and a decision table.

What DeepSeek shipped on September 10

DeepSeek’s changelog calls V4.1 Flash “the smallest model in our new architecture family, with native multimodal visual understanding.” Three things matter for anyone currently paying for K3:

  1. A new model ID. The canonical name is now deepseek-flash. The legacy names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work, but requests are “served by the DeepSeek-V4.1-Flash model and billed at the Flash price,” per the pricing page. If your code still says deepseek-v4-flash, you already switched to V4.1.
  2. Native vision and lower prices. Images go straight into the main model (the August 21 vision-exp SKU is retired), and peak rates fell to $0.30 input / $1.20 output.
  3. A reversal on V4 Pro. Launch-day coverage, including SiliconANGLE, reported that deepseek-v4-pro calls would route to V4.1 Flash from September 14. DeepSeek’s changelog now says it will “continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.” The pricing page still lists deepseek-v4-pro as DeepSeek-V4-Pro-0813.

The weights are on Hugging Face under the MIT license, along with a technical report. SCMP’s coverage led with the angle Kimi users care about: DeepSeek says the new Flash beats Kimi K3 on cyber and coding benchmarks.

Kimi K3 vs DeepSeek V4.1 Flash: specs side by side

Kimi K3DeepSeek V4.1 Flash
ReleasedJuly 16, 2026 (weights July 27)September 10, 2026
API model IDkimi-k3deepseek-flash
ArchitectureMoE, 2.8T total, 104B activeCausal encoder-decoder MoE, 552B backbone + 196B Engram memory
Active parameters104B per token8B during prefill, 16B during decode
Context window1,048,576 tokens1M tokens
Max output131,072 default, 1,048,576 max384K max
VisionNative; images plus video via file IDsNative; images only (base64, URL or file ID)
ThinkingAlways on; low / high / max, default maxOn by default (high); can be disabled; low / high / max
API formatsOpenAI-compatibleOpenAI, Anthropic and Responses API formats
Weights on disk~1.56TB (MXFP4)~510GB (mixed INT8 / FP8 / BF16)
LicenseKimi K3 License (MIT-style with conditions)MIT
Output speed (Artificial Analysis)35.2 tok/s221.7 tok/s
Time to first token (AA)3.81s0.92s
AA Intelligence Index44 (#3 of 114)39 (#7 of 114)

Sources: Kimi’s K3 quickstart and pricing page, DeepSeek’s pricing page, thinking-mode guide and vision guide, the V4.1 Flash model card, and Artificial Analysis pages for Kimi K3 and V4.1 Flash, checked September 25. Weight sizes were measured from the Hugging Face file listings. Some third-party write-ups cite 40 for Flash on the AA index; the live page showed 39 when we checked.

Active parameters explain most of the price and speed gap. Flash’s decoder builds its KV cache from the encoder’s final hidden states, so only 8B parameters fire per token while reading a prompt, which is why DeepSeek pitches it for “input-heavy agentic workloads.” K3 activates 104B on every token, 6.5 to 13 times more.

Reasoning controls differ too. The model card describes a 1–100 effort dial, but the API thinking-mode guide only documents low, high and max, so treat the numeric dial as self-hosting only for now. Flash can also turn thinking off entirely. K3 can’t: Kimi’s docs say “K3 always has thinking mode enabled.”

For the full K3 architecture, see our Kimi K3 specifications page.

V4.1 Flash vs V4-Flash-0731: what actually changed

A lot of comparisons from early September, including several that pitted K3 against DeepSeek’s July Flash, are already outdated. Here is what separates the two DeepSeek models:

V4-Flash-0731V4.1 Flash
ReleasedJuly 31, 2026September 10, 2026
What it isRe-post-trained V4-Flash; same architecture and size as the previewNew architecture, trained from scratch on 45T multimodal tokens
Total / active params284B / 13B552B backbone + 196B Engram / 8B–16B
VisionNone (separate vision-exp SKU from Aug 21)Native
KV cacheBaseline~1/4 of V4-Flash (890 bytes per token, FP4)
Context / max output1M / 384K1M / 384K
Weights on Hugging Face~167GB~510GB
LicenseMITMIT
Peak price (hit / miss / output)$0.014 / $0.44 / $1.32$0.006 / $0.30 / $1.20
API status todayRetired; name routes to V4.1 FlashCurrent (deepseek-flash)
Terminal-Bench 2.182.790.6
DeepSWE54.474.2

0731 benchmarks are from DeepSeek’s July 31 changelog entry; V4.1 figures from the model card. The 0731 prices were in effect August 16 to September 10, as listed by Codersera’s V4 pricing guide and matching the August pricing update in DeepSeek’s changelog.

In short, 0731 was the cheap text-only model; V4.1 Flash is a different model that is bigger in total, smaller in active parameters, multimodal and cheaper to call. 0731 only wins on download size.

Pricing: the sticker numbers

Kimi K3 bills a flat rate at any hour. DeepSeek charges double during peak hours: 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours, weekends included, are off-peak at half price.

Per 1M tokensKimi K3V4.1 Flash (peak)V4.1 Flash (off-peak)V4-Flash-0731 (peak, retired)
Input, cache hit$0.30$0.006$0.003$0.014
Input, cache miss$3.00$0.30$0.15$0.44
Output$15.00$1.20$0.60$1.32
Cache write$3.00 (5-min TTL), $6.00 (1-hour TTL)not billed separatelynot billed separatelynot billed separately
Concurrency limitby top-up tier2,5002,500n/a

One detail most comparisons miss: Kimi’s pricing page now lists two cache-write tiers for K3. The default 5-minute TTL costs the same $3.00 as normal input; the optional 1-hour TTL costs $6.00 on the write and only pays off for heavily reused prefixes.

On paper, Flash is 10x cheaper on uncached input and 12.5x cheaper on output at peak, and 20x and 25x cheaper off-peak. The cache-hit gap is 50x to 100x.

Timezone matters. Peak windows cover only 35 of the week’s 168 hours, all between 9pm and 6am US Eastern, so a US team working normal hours almost always pays off-peak rates. European mornings (08:00–12:00 CEST) and Indian late mornings (11:30–15:30 IST) are peak.

You may also see Flash quoted at “$0.14 / $0.28.” That was DeepSeek’s V4-Flash list price before the peak/off-peak system started on August 16. It is not a current official rate.

Advertised vs real cost

List-price ratios overstate how much you will save, for one reason: Flash writes more tokens. Artificial Analysis recorded 250M output tokens for V4.1 Flash (max effort) across its Intelligence Index, against 160M for K3 (max), about 1.56x more. Its per-task cost figures reflect that:

  • Kimi K3: $2.00 per Intelligence Index task
  • V4.1 Flash: $0.27 per task

That is a 7.4x real-world gap, not the 12.5x to 25x suggested by output prices. It is still large, and Flash also finished its work much faster (221.7 vs 35.2 output tokens per second in AA’s measurements).

Three workloads at official list prices, assuming identical token counts (so treat them as upper bounds on Flash’s advantage):

WorkloadKimi K3V4.1 Flash peakV4.1 Flash off-peak
One uncached call: 100K in, 20K out$0.60$0.054$0.027
Coding-agent session: 50 turns, each 200K cached + 5K new input + 3K output$6.00$0.32$0.16
Monthly: 1B input at 90% cache hit + 50M output$1,320$95.40$47.70

The agent-session row is where K3 hurts most: even at a 97% cache-hit rate, its $0.30 cached price and $15 output price dominate the bill, while Flash’s 10M cached tokens cost six cents. Two levers change the math. Flash can run with thinking off for simple steps, which K3 cannot. And cache-hit rate matters on both sides, so measure it on your own traffic.

Benchmarks: what’s verified and what’s vendor-reported

The only head-to-head table that includes both models comes from DeepSeek itself, in the V4.1 Flash model card. All numbers are at maximum reasoning effort, and DeepSeek ran the evaluations in its own harness.

BenchmarkKimi K3V4.1 FlashV4-Flash (0731)
GPQA Diamond92.990.989.9
HLE (no tools)43.536.837.8*
MathArena Apex65.665.658.6
Terminal-Bench 2.188.390.682.7
Terminal-Bench 3.017.730.07.6
Terminal-Bench 4.012.631.27.0
DeepSWE v1.167.574.254.4
NL2Repo-Bench58.064.054.2
CyberGym80.088.176.7
HLE with tools59.863.951.5
AutomationBench46.754.837.7
Agent’s Last Exam27.631.825.2
Chartography (with tools)68.178.9—
BabyVision (with tools)85.789.6—
ZeroBench-main (with tools, pass@5)41.049.0—

*Text-only HLE subset. Source: DeepSeek V4.1 Flash model card, “Comparison with frontier models (Max reasoning effort).”

How to read this honestly:

  • K3’s Terminal-Bench 2.1 (88.3) and DeepSWE (67.5) match Moonshot’s launch numbers. GPQA differs slightly: 92.9 here vs Moonshot’s 93.5.
  • The biggest gaps can’t be cross-checked. On Terminal-Bench 3.0 and 4.0 (K3 17.7 and 12.6, Flash 30.0 and 31.2) we found no Moonshot-published scores. Treat them as DeepSeek’s measurements.
  • The harness moves the headline. DeepSeek’s scaffold table shows Flash’s DeepSWE score ranges from 65.5 (OpenCode) to 74.2 (mini-SWE), with 69.8 in Claude Code and 65.6 in Codex, which is below K3’s 67.5. The 74.2 headline is the best case. (DeepSeek’s changelog also lists NL2Repo 65.4 where the model card says 64.0.)
  • K3 keeps its reasoning lead, on GPQA and HLE here and on the independent Artificial Analysis index (44 vs 39).

Net: Flash is at least competitive with K3 on agentic coding, probably better in DeepSeek’s preferred scaffolds, and far ahead on value. K3 is still the stronger generalist. No independent same-harness coding head-to-head exists yet.

Licenses: MIT vs the Kimi K3 License

Both models ship open weights, but the terms differ.

DeepSeek V4.1 Flash: MIT. Keep the copyright notice and you can use, modify, host, resell and fine-tune it. There are no revenue thresholds or attribution rules. The same was true of 0731.

Kimi K3: the Kimi K3 License. The license file on Hugging Face starts from MIT wording and adds two conditions:

  1. Model-as-a-Service clause. If you run an API business that gives third parties meaningful control over inference or fine-tuning, and your group’s revenue exceeds $20 million over any 12 months, you need a separate agreement with Moonshot before commercial use. Embedding K3 inside a product feature, or simply relaying requests to another host, does not count as MaaS.
  2. Attribution clause. Commercial products with more than 100 million monthly active users or $20 million in monthly revenue must display “Kimi K3” prominently in the UI.

Neither clause applies to internal use or to use through Moonshot’s official products or certified inference partners. For most builders the difference is academic; for an inference provider or a huge consumer app, MIT is simpler. We go deeper in our Kimi K3 license explainer.

Self-hosting: 510GB vs 1.56TB

Self-hosting is where Flash’s smaller footprint pays off most directly.

  • Kimi K3 is about 1.56TB of MXFP4 weights and needs roughly 1.7–1.84TB of GPU memory to serve. Working configs start at 8×B300 or 16×H200, and 8×H100 is impossible. Full math in our VRAM requirements guide.
  • V4.1 Flash is about 510GB on Hugging Face (measured from the repo’s file listing), and its KV cache is tiny at 890 bytes per token. On paper that fits one 8×H200 node (1,128GB) with room for long contexts, but DeepSeek publishes no reference hardware config, so treat it as arithmetic, not a tested deployment.

One snag: the release has no Jinja chat template. DeepSeek ships a Python reference encoder and a Rust toolkit, deepseek-recipe, for prompt formatting, so expect some lag before every serving stack handles the new architecture cleanly.

When to switch from K3 to DeepSeek Flash (and when not to)

Your situationRecommendation
High-volume coding or automation agents, cost is the main constraintSwitch most traffic to V4.1 Flash. Keep K3 as an escalation model for tasks Flash fails
Latency-sensitive interactive productV4.1 Flash: 0.92s TTFT and ~220 tok/s vs 3.81s and ~35 tok/s
Simple classification, extraction or short repliesV4.1 Flash with thinking off or low. K3 can’t turn thinking off
Hard reasoning, science, research questionsStay on K3. It leads GPQA and HLE and the AA index
Outputs longer than 384K tokensK3 (1M max output)
Video inputK3. DeepSeek’s vision guide covers images only
Screenshot or chart-heavy agentsTest both. Flash scores higher on DeepSeek’s visual-agent benchmarks; verify on your images
You run a public inference API or need zero license conditionsV4.1 Flash (MIT)
Self-hosting on one nodeV4.1 Flash (~510GB) over K3 (~1.56TB)
Already built on Kimi Code and tuned to K3’s behaviorStay, and A/B Flash on a slice of traffic before moving

If you were on V4-Flash-0731 and comparing it with K3: the comparison you started no longer exists on DeepSeek’s API. Your deepseek-v4-flash calls now hit V4.1 Flash. Re-run your evals, because behavior changed even though the model name did not.

A sensible setup is Flash by default and K3 for failures or flagged hard tasks. Both expose OpenAI-compatible endpoints, so it is a model-ID switch, not a rewrite. For Zhipu’s cheap challenger, see Kimi K3 vs GLM-5.3 Flash.

Caveats before you migrate

  • Benchmarks are self-reported. The only table with both models is DeepSeek’s. Budget on cost per task (about 7.4x per AA), not per-token prices.
  • Legacy names are “temporarily” routed. DeepSeek’s changelog uses that word. Move your code to deepseek-flash now rather than waiting for the alias to break.
  • Data jurisdiction. Both first-party APIs are run by Chinese companies. If that matters to your legal team, look for US or EU third-party hosts for either model and check their region and retention terms. For K3, those options are mapped in our providers guide.
  • Prices move. DeepSeek changed Flash pricing twice in four weeks (August 16 and September 10). Re-check before you commit a budget.

FAQ

Is DeepSeek V4.1 Flash better than Kimi K3? On the agentic coding and automation benchmarks DeepSeek published, yes: 90.6 vs 88.3 on Terminal-Bench 2.1, 74.2 vs 67.5 on DeepSWE v1.1 and 88.1 vs 80.0 on CyberGym. But those are DeepSeek’s own runs in its own harness. On independent scoring, Artificial Analysis still ranks Kimi K3 higher (Intelligence Index 44 vs 39), and K3 leads on GPQA Diamond and Humanity’s Last Exam even in DeepSeek’s table. Flash is the better value; K3 is still the deeper model.

How much cheaper is DeepSeek V4.1 Flash than Kimi K3? On list price, Flash is 10x cheaper on uncached input and 12.5x cheaper on output at DeepSeek’s peak rates, and 20x and 25x cheaper off-peak. Cached input is 50-100x cheaper. Real per-task savings are smaller because Flash writes more tokens: Artificial Analysis measured $0.27 per Intelligence Index task for Flash vs $2.00 for K3, about 7.4x.

What changed from DeepSeek V4-Flash-0731 to V4.1 Flash? Almost everything. V4-Flash-0731 was a re-post-trained version of the 284B/13B-active V4-Flash, text-only. V4.1 Flash is a new causal encoder-decoder architecture with a 552B backbone plus a 196B Engram memory, 8B active parameters during prefill and 16B during decode, native image input, and a KV cache about a quarter the size. It is also cheaper: $0.30/$1.20 peak vs $0.44/$1.32 for 0731.

Can I still use DeepSeek V4-Flash-0731 through the API? Not really. DeepSeek retired V4 Flash and V4 Flash Vision Exp on September 10, 2026. The old model names deepseek-v4-flash and deepseek-v4-flash-vision-exp still work, but they are served by V4.1 Flash and billed at the Flash price. If you need 0731’s exact behavior, the MIT-licensed weights are still on Hugging Face for self-hosting.

Does DeepSeek V4.1 Flash support images like Kimi K3? Yes. Unlike 0731, V4.1 Flash has native image understanding in the main model, accessed through the deepseek-flash model ID. DeepSeek’s API accepts base64, HTTPS URLs or file IDs, up to 600 images per request, and caps each image at 1,024 tokens. Kimi K3 also has native vision and additionally accepts video through uploaded file IDs, but does not accept public image URLs.

What are DeepSeek’s peak hours? Peak pricing applies from 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday through Friday, excluding Chinese public holidays. Everything else, including weekends, is billed at half price. US business hours fall entirely in the off-peak window; European mornings and Indian late mornings do not.

Which license is more permissive, MIT or the Kimi K3 License? MIT. DeepSeek V4.1 Flash weights are plain MIT with no usage conditions beyond keeping the notice. The Kimi K3 License is MIT-style but adds two conditions: API resellers with over $20M in revenue over 12 months need a separate agreement with Moonshot, and products above 100M monthly users or $20M monthly revenue must display Kimi K3 in the UI.

Should I move my coding agent from Kimi K3 to DeepSeek V4.1 Flash? Run a side-by-side on your own tasks first, but for most teams the answer is to route the bulk of routine agent work to Flash and keep K3 for the hardest tasks. Flash is roughly six times faster in Artificial Analysis testing and far cheaper. Stay on K3 if your tasks depend on long outputs over 384K tokens, video input, or the broader knowledge that shows up in K3’s GPQA and HLE leads.

The bottom line

DeepSeek V4.1 Flash is the strongest value case against Kimi K3 so far. It matches or beats K3 on most agentic benchmarks DeepSeek ran, costs about a seventh as much per real task, runs about six times faster and ships under plain MIT. K3 keeps stronger general reasoning (confirmed independently), a 1M-token output ceiling, video input and the Kimi ecosystem. For most teams the move is to demote K3, not abandon it: Flash by default, K3 for escalation, and your own evals, because neither vendor’s table ran your workload.

Keep reading: Kimi K3 API providers, Kimi K3 VRAM requirements and the open weights breakdown.