← All guides
Kimi K3 vs Claude Fable 5: Benchmarks & Verdict (2026)

Kimi K3 vs Claude Fable 5: Benchmarks & Verdict (2026)

July 19, 2026 · kimi k3 vs claude fable 5 · kimi k3 vs claude · fable 5 vs kimi k3 · ai model comparison

Six weeks after Anthropic launched Claude Fable 5 (June 9, 2026) as its new Mythos-class flagship, Moonshot AI answered with Kimi K3 (July 16, 2026) — a 2.8-trillion-parameter open-weight model that beat Fable 5 on every single agentic benchmark in its launch suite, at one-third the API price.

But the real story is more nuanced than “K3 wins” — and more interesting. Buried in Moonshot’s own footnotes is a detail almost nobody is talking about: Fable 5’s safety routing kicked in on 35% of tasks in one key benchmark, meaning part of its score was actually earned by the weaker Claude Opus 4.8. This comparison covers all of it — every benchmark number, exact token pricing, independent test results, the caveats on both sides, and a clear verdict for every type of user.

The two models at a glance

Kimi K3Claude Fable 5
MakerMoonshot AIAnthropic
ReleasedJuly 16, 2026June 9, 2026
Model classOpen-weight flagshipMythos-class (closed flagship)
Parameters2.8T MoE (16/896 experts, ~50B active)Undisclosed
Context window1M tokens1M tokens
Max outputNot capped publicly128K tokens
Thinking modeMax effort (default at launch)Adaptive thinking (default)
Input price (per 1M)$3.00 ($0.30 cached)$10.00 ($1.00 cached)
Output price (per 1M)$15.00$50.00
Open weightsYes — July 27, Modified MITNo
Knowledge cutoffJanuary 31, 2026
Output speed~62 tok/s~57 tok/s

Pricing: the 3.3× gap, tier by tier

Most comparisons quote a single price. Here’s the full rate card — the gap holds on every billing tier:

Billing tier (per 1M tokens)Kimi K3Claude Fable 5K3 cheaper by
Input (standard)$3.00$10.003.3×
Input (cache hit/read)$0.30$1.003.3×
Output$15.00$50.003.3×
Cache write$12.50 (5 min) / $20 (1 hr)
Batch discount50% off

In Moonshot’s benchmark suite, the average cost per completed task came out to $0.94 for K3 vs $2.75 for Fable 5 — almost exactly the 3× you’d predict from token prices. K3’s economics lean on Mooncake disaggregated inference with a >90% cache-hit rate in coding workloads, meaning most real-world input bills land at the $0.30 tier.

Two honest counterpoints: Fable 5 offers a 50% batch discount (K3 doesn’t publish one yet), which narrows the gap to ~1.7× for offline workloads. And Fable 5’s adaptive thinking can save money by reasoning less on easy tasks — K3 launched locked at max thinking effort, with low/high modes still to come.

The official benchmarks: K3 wins all six

Moonshot’s launch suite tested every model at maximum thinking effort (temperature 1.0, top-p 1.0), each under its designated agentic harness — KimiCode, Claude Code, or Codex. Here’s every result, K3 vs Fable 5:

BenchmarkKimi K3Claude Fable 5Margin
Terminal-Bench 2.188.384.6K3 +3.7
Program Bench77.876.8K3 +1.0
SWE Marathon42.035.0K3 +7.0
BrowseComp91.288.0K3 +3.2
SpreadsheetBench 234.834.7K3 +0.1
Automation Bench30.829.1K3 +1.7

Official Kimi K3 coding benchmarks vs Claude Fable 5, GPT-5.6 Sol and Claude Opus 4.8 — DeepSWE, Terminal-Bench 2.1, Program Bench, SWE Marathon

Official Kimi K3 agentic benchmarks vs Claude Fable 5 — BrowseComp, SpreadsheetBench 2, Automation Bench, GDPval-AA

Now the part that makes this comparison worth reading — the caveats buried in the footnotes, which cut both ways:

The 35% fallback issue (hurts Fable 5’s score)

For SWE Marathon — K3’s biggest win, 42.0 vs 35.0 — Moonshot’s footnote states: “Claude Fable 5 hit fallbacks on 35% of the tasks in our evaluation, which may have negatively impacted its measured performance.”

What’s a “fallback”? Anthropic ships Fable 5 with safety classifiers that route cybersecurity, biology, chemistry, and model-distillation-related requests to Claude Opus 4.8 instead of answering directly. (The unrestricted version, Claude Mythos 5, is locked behind Anthropic’s Project Glasswing trusted-access program.) So on more than a third of SWE Marathon tasks, the model doing the work was effectively Opus 4.8 — which scores 40.0 on its own. Fable 5’s true ceiling on this benchmark is almost certainly higher than 35.0, though how much higher is impossible to say from public data. It’s a fair caveat — but it’s also a real-world limitation: if you run coding agents that touch security-sensitive code, this routing will affect you too, not just benchmark scores.

The context-compaction caveat (hurts K3’s score)

For BrowseComp, Moonshot used the context-compaction strategy from Anthropic’s model cards, triggered at 300K tokens — a fraction of K3’s 1M window. Run with the full window and no compaction, K3 scores 90.4–91.2 without management tricks. In other words, K3’s published 91.2 was achieved with one hand tied.

The harness and hardware notes

K3 was evaluated under its own KimiCode harness on several benchmarks while Fable 5 used Claude Code or Terminus 2 — harness choice affects scores. Some PostTrain Bench runs used H20 GPUs instead of the official H100 setting. And KCB 2.0 is an in-house Moonshot benchmark. None of this invalidates the results, but vendor-run benchmarks always deserve independent confirmation — which brings us to the third-party data.

Independent testing: what non-vendor sources say

This is where an honest comparison has to widen the lens:

LMArena (community blind voting) — the strongest independent signal for K3:

  • Frontend Code Arena: K3 is #1 with 1679 Elo vs Fable 5’s 1631 — a 76% pairwise win rate vs Fable 5’s 63%
  • K3 is #1 in 6 of 7 frontend domains (Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Simulations, Content Creation Tools)
  • The one domain K3 loses? Gaming — where Fable 5 holds #1
  • On the general Text Arena, K3 ranks #9 (up from K2.6’s #38) — solid, but not beating the closed flagships there

Anthropic-reported and aggregator benchmarks for Fable 5 (no public K3 equivalents yet):

  • SWE-bench Verified: 95% and SWE-bench Pro: 80.3% — ahead of Opus 4.8 (69.2%), GPT-5.5 (58.6%), and Gemini 3.1 Pro (54.2%)
  • GPQA Diamond: 92.6% for graduate-level scientific reasoning

Artificial Analysis places K3 at #4 of 189 models (score 57) — the best result ever for an open-weight model — but still behind Fable 5 and GPT-5.6 Sol on its aggregate intelligence index. Moonshot itself says the same thing in its launch post: K3 “still trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol” in overall user experience.

The honest synthesis: on agentic, tool-using, production-style tasks, K3 matches or beats Fable 5 nearly everywhere it’s been measured. On general knowledge chat and some academic reasoning suites, Fable 5’s closed-model polish still shows.

Speed: neither is a sprinter

Independent measurements put K3 at roughly 62 output tokens/sec with ~2s time-to-first-token, and Fable 5 at roughly 57 tok/s. Both are below the median for their price tiers — these are thinking models tuned for long autonomous runs, not snappy chat. If your product needs instant conversational responses, both models will feel deliberate; K3’s coming low-effort mode and Fable 5’s adaptive thinking are the respective answers.

Context: both 1M tokens — but not equal economics

On paper it’s a tie: both models offer a 1M-token context window. The difference is cost and speed at scale. Filling a 1M window costs $3 with K3 (or $0.30 on cache hits) vs $10 with Fable 5 — and K3’s Kimi Delta Attention architecture decodes long contexts up to 6.3× faster than its predecessor, with a prefill-cache implementation contributed to vLLM. Fable 5 counters with a larger 128K max output and mature prompt-caching. For whole-codebase or book-length workloads, K3 does the same job for a third of the price.

Openness: the fundamental divide

On July 27, 2026, K3’s full weights release under a Modified MIT license (MXFP4-quantized, with MXFP8 activations) alongside a technical report — the largest open-weight model ever. Moonshot recommends supernode deployments (64+ accelerators) for self-hosting, so this matters to companies and labs, not hobbyists — but for them it unlocks what Fable 5 can never offer: self-hosting for data residency, deep fine-tuning, and zero vendor lock-in.

Fable 5’s closed model has its own upside: zero infrastructure decisions, availability across the Claude API, Amazon Bedrock, Azure Foundry, Google Vertex, GitHub Copilot, and a 319-page system card with thousands of automated behavioral audits — the most extensively safety-documented frontier model release to date. Anthropic also offers fine-tuning on its platform. For regulated enterprises that want a vendor to carry the safety burden, that matters.

Real-world capability: beyond benchmarks

K3’s launch demos were unusually concrete: it wrote MiniTriton, a working Triton-like GPU compiler from scratch (tile-level IR over MLIR, PTX codegen, beating Triton on some workloads); designed a functioning chip in a 48-hour autonomous run (1.46M standard cells, timing closed at 100 MHz on Nangate 45nm); reproduced astrophysics research (the I–Love–Q universal relations) in 2 hours vs the typical 1–2 weeks, reviewing 20+ papers and generating 3,000+ lines of Python; and edited its own teaser video from 56 source clips. In Moonshot’s kernel-optimization test, K3 “performed competitively with Fable 5 (with fallback) and substantially outperformed Opus 4.8, GPT-5.6 Sol, and GPT-5.5.”

Fable 5’s strengths show up differently: Anthropic positions it for the hardest long-horizon knowledge work — codebase-wide migrations, dense document synthesis, multi-day agentic runs — with adaptive thinking deciding how hard to reason per request. Its SWE-bench Pro 80.3% remains the strongest published score on that academic suite.

Kimi K3 vs Claude Fable 5: which should you choose?

Your situationPick
Agentic coding at scale (SWE Marathon +7 pts, 3× cheaper)Kimi K3
Frontend / UI development (#1 Arena, 76% win rate)Kimi K3
Game developmentClaude Fable 5 (#1 Gaming domain)
1M-context workloads on a budgetKimi K3
Open weights, self-hosting, fine-tuning your own stackKimi K3
Academic reasoning suites (SWE-bench Pro, GPQA)Claude Fable 5
General chat polish and UXClaude Fable 5 (narrowly)
Enterprise safety documentation & complianceClaude Fable 5
Offline/batch workloads (50% batch discount)Claude Fable 5 (gap narrows to ~1.7×)

The smartest move costs nothing: test both on your actual workload. K3 is free to try in the Kimi app — no credit card, and an invite link gets you bonus membership credits. Take your 10 hardest real tasks and run them through both. Here’s how to get K3 on your PC.

FAQ

Is Kimi K3 better than Claude Fable 5? On agentic and coding benchmarks, yes — K3 wins all six head-to-head tests in Moonshot’s launch suite and leads the Frontend Code Arena (1679 vs 1631 Elo). On aggregate intelligence rankings and general-chat polish, Fable 5 still leads. At one-third the price, K3 offers better value for most production workloads.

Why did Fable 5 score only 35.0 on SWE Marathon? Partly because of Anthropic’s safety routing: Moonshot’s footnotes report Fable 5 hit fallbacks on 35% of tasks, where requests were handed to the weaker Claude Opus 4.8. Its unrestricted ceiling is likely higher — but the routing affects real-world security-adjacent coding too, not just benchmarks.

Is Kimi K3 cheaper than Claude Fable 5? Yes — exactly 3.3× cheaper on every token tier ($3/$15 vs $10/$50 per 1M tokens; $0.30 vs $1.00 cached input). Measured cost per benchmark task: $0.94 vs $2.75. Fable 5’s 50% batch discount narrows the gap for offline workloads.

Kimi K3 vs Claude Fable 5 for coding — which wins? K3 for most developers: it sweeps the coding benchmarks (Terminal-Bench, Program Bench, SWE Marathon), ranks #1 on the Frontend Code Arena, and costs a third as much. Fable 5 holds the edge in game development and on academic suites like SWE-bench Pro (80.3%).

Can Kimi K3 replace Claude in my company? For agentic, coding, research and office-automation workloads — likely yes, at a third of the cost, with open weights for self-hosting after July 27. If you need Anthropic’s compliance documentation, GitHub Copilot integration, or top academic-reasoning scores, Fable 5 keeps its edge. Test both on real tasks before migrating.

Which is faster, K3 or Fable 5? Roughly tied: K3 outputs ~62 tok/s vs Fable 5’s ~57 tok/s. Both are deliberate “thinking” models built for long autonomous runs rather than instant chat responses.

The verdict

Kimi K3 vs Claude Fable 5 is the first genuinely competitive open-vs-closed matchup at the frontier. K3 sweeps the agentic benchmarks, leads the community-voted coding arena, matches the 1M context window, and costs exactly 3.3× less on every pricing tier — with open weights landing July 27. Fable 5 answers with stronger academic-reasoning scores, general-chat polish, gaming, a 319-page safety dossier, and the deepest enterprise ecosystem in AI.

If you’re spending real money on agents or code in 2026, K3 is the rational default. If you need the last few points of reasoning quality — and a vendor to sign the compliance paperwork — Fable 5 earns its premium. For the full four-way picture including GPT-5.6 Sol and Opus 4.8, see our frontier model comparison, or start from the complete Kimi K3 guide.

Try it yourself

Sign up for Kimi through our invite link and both of us get free bonus membership credits — up to a full year, at no cost to you.

Claim Free Credits →