Kimi K3 vs Claude Opus 5: Benchmarks, Pricing & Which Wins
July 25, 2026 · kimi k3 vs claude opus 5 · claude opus 5 vs kimi k3 · kimi k3 vs claude · claude opus 5 · ai model comparison
Eight days. That’s all the time Moonshot AI’s Kimi K3 (July 16, 2026) had as the newest frontier model on the block before Anthropic answered with Claude Opus 5 (July 24, 2026) — a “thoughtful and proactive” model that Anthropic claims delivers near-Fable 5 intelligence at half the price.
This is the matchup the industry actually cares about. Not K3 vs the $10/$50 Mythos-class Fable 5 — but K3 vs Anthropic’s mainstream flagship, the new default on Claude Max, priced in the same ballpark and aimed at the same agentic, coding-heavy workloads. And the independent data is already in: K3 wins 3 of 5 head-to-head benchmarks, costs 1.7× less per token, and ships open weights in two days. Opus 5 fires back with state-of-the-art knowledge work, a five-level effort ladder, and the deepest enterprise distribution in AI.
Here’s every number that matters — pricing, benchmarks, context, speed, openness — plus a clear verdict for every type of user.
The two models at a glance
| Kimi K3 | Claude Opus 5 | |
|---|---|---|
| Maker | Moonshot AI | Anthropic |
| Released | July 16, 2026 | July 24, 2026 |
| Model class | Open-weight frontier flagship | Closed frontier (below Mythos tier) |
| Parameters | 2.8T MoE (16/896 experts active) | Undisclosed |
| Context window | 1M tokens (1,048,576) | 1M tokens (1,000,000) |
| Max output | Not publicly capped (up to 1M listed) | 128K tokens |
| Thinking mode | Always on; low / high / max effort | On by default; off allowed at effort ≤ high |
| Effort control | low / high / max; plus k3-256k variant (~half quota) | Full ladder: low / medium / high / xhigh / max |
| Input price (per 1M) | $3.00 ($0.30 cached) | $5.00 ($0.50 cached) |
| Output price (per 1M) | $15.00 | $25.00 |
| Modalities | Text, image, video | Text, image, audio, video |
| Open weights | Yes — July 27, Modified MIT | No |
| Availability | Kimi app, Kimi Work, Kimi Code, API | Claude API, Bedrock, Google Cloud, Foundry; default on Claude Max |
| Output speed | ~62 tok/s, ~2s TTFT | Effort-dependent |
One positioning note that matters: Opus 5 is not Anthropic’s top model. The Mythos tier (Fable 5 / Mythos 5) still sits above it. Anthropic’s pitch for Opus 5 is efficiency — “close to the frontier intelligence of Claude Fable 5 at half the price.” K3’s pitch is nearly identical, except the baseline it’s measured against is everyone.
Pricing: K3 is 1.7× cheaper on every tier
| Billing tier (per 1M tokens) | Kimi K3 | Claude Opus 5 | K3 cheaper by |
|---|---|---|---|
| Input (standard) | $3.00 | $5.00 | 1.67× |
| Input (cache hit/read) | $0.30 | $0.50 | 1.67× |
| Cache write (5-min TTL) | — (automatic via Mooncake) | $6.25 | — |
| Cache write (1-hour TTL) | — (automatic via Mooncake) | $10.00 | — |
| Output | $15.00 | $25.00 | 1.67× |

How cache billing actually works on each side: Anthropic charges explicitly for cache writes — 1.25× the input rate ($6.25/MTok) for a 5-minute TTL and 2× ($10/MTok) for a 1-hour TTL — then $0.50/MTok on cache reads. Moonshot’s Mooncake disaggregated inference handles caching automatically with no published write charge: reads bill at $0.30/MTok, and Moonshot reports >90% cache-hit rates on coding workloads, so most real-world K3 input bills land at the cached tier. For long agentic loops that re-read the same context every turn, that automatic caching is a structural cost advantage Opus 5’s manual cache writes can’t fully match.
On a representative agentic workload mix (7:2:1 cached:input:output), that lands at roughly $2.31 per million blended tokens for K3 vs $3.71 for Opus 5 — about 60% of the cost. It’s not the dramatic 3.3× gap from the K3 vs Fable 5 comparison, but it compounds brutally at scale: a team spending $50K/month on Opus 5 tokens is looking at roughly $20K in savings on K3 for the same volume.
Two honest counterpoints. First, Opus 5’s effort ladder is a real cost feature: at low or medium effort it burns far fewer reasoning tokens on easy tasks, and Anthropic’s launch charts show Opus 5 delivering the best performance-per-dollar of any Anthropic model at every effort setting. K3 has since shipped low/high effort modes and a k3-256k variant that halves quota for ≤256K workloads — narrowing this advantage, though Opus 5’s five-level ladder is still more granular. Second, Anthropic offers a 50% Batch API discount for offline workloads; Moonshot doesn’t publish one yet.
Head-to-head benchmarks: the independent data
Within hours of Opus 5’s launch, independent aggregators had both models on the same evals. Here’s the actual head-to-head — no vendor picking the test suite:
| Benchmark | Kimi K3 | Claude Opus 5 | Winner |
|---|---|---|---|
| GDPval-AA (knowledge work) | 55.6 | 62.0 | Opus 5 +6.4 |
| Humanity’s Last Exam (reasoning) | 56.0 | 64.7 | Opus 5 +8.7 |
| AutomationBench (Zapier, business tasks) | 30.8 | 26.0 | K3 +4.8 |
| BrowseComp (agentic web research) | 91.2 | 90.8 | K3 +0.4 |
| DeepSWE 1.1 (agentic coding) | 69.0 | 68.8 | K3 +0.2 (tie) |

The pattern is the same one we saw against Fable 5: K3 wins the agentic, tool-using, long-horizon categories; Claude wins knowledge work and academic reasoning. Except this time the price gap is 1.7× instead of 3.3× — and Opus 5’s reasoning wins are larger than its agentic losses.
The AutomationBench discrepancy worth understanding
Anthropic’s launch post claims Opus 5’s AutomationBench pass rate is “around 1.5× the next-best model for the same cost per task,” and that even at its lowest effort setting it passes more tasks than any other model. The independent raw scores say K3 30.8, Opus 5 26.0. Both can be true: Anthropic is measuring cost-adjusted throughput across effort settings, while the raw number is single-config pass rate. If you’re benchmarking for a purchase decision, measure cost per completed task on your own workload — that single metric cuts through both framings.
Anthropic’s own evals (not yet independently verified)
Anthropic also published vendor-run results where K3 wasn’t tested head-to-head, so treat these as claims awaiting confirmation:
- Frontier-Bench v0.1: new state-of-the-art, more than doubling Opus 4.8 at a lower cost per task
- CursorBench 3.2: within 0.5% of Fable 5’s peak at max effort, at half the cost per task
- ARC-AGI 3: three times the score of the next-best model on novel problem solving
- OSWorld 2.0 (computer use): surpasses Fable 5’s best result at just over a third of the cost
- Cybersecurity: still behind Mythos 5 — Anthropic’s own safety positioning keeps the hardest cyber capability in the trusted-access tier
K3’s independent credentials, meanwhile, are already battle-tested: #1 on LMArena’s Frontend Code Arena (1,679 Elo) — ahead of Fable 5, GPT-5.6 Sol, and everything else — 88.3 on Terminal-Bench 2.1, 42.0 on SWE Marathon, and #4 of 189 models on Artificial Analysis (score 57), the best result ever for an open-weight model. Aggregate scores for Opus 5 are still landing as evaluators complete their runs.
Where Claude Opus 5 wins
Knowledge work and hard reasoning. GDPval-AA 62.0 and HLE 64.7 are the strongest scores in this price class, and they’re not close — 6 to 9 points over K3. If your workload is “read these 40 documents and produce a correct analysis,” Opus 5 is the better engine today.
Granular cost and latency control. Five effort levels (low → max), thinking that can be disabled at high and below, and per-request tuning of the speed/intelligence tradeoff. K3 now ships low/high/max effort modes too (added post-launch), but three settings versus five — and no option to disable thinking — leaves Opus 5 with the finer dials. For high-volume products where most requests are easy, Opus 5’s low-effort tier can undercut K3’s effective cost on simple calls.
Enterprise distribution and trust. Available day one on the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry; the new default on Claude Max; backed by Anthropic’s compliance documentation, safety routing, and fallback infrastructure (including a new server-side "default" fallback mode and mid-conversation tool changes, both in beta). If procurement needs a US vendor with audit paperwork, this is the end of the discussion.
Larger guaranteed output. A hard 128K max output ceiling beats K3’s publicly uncapped-but-unverified limit for teams that generate long artifacts in one pass.
Where Kimi K3 wins
Agentic coding and automation. DeepSWE 1.1, AutomationBench, BrowseComp — K3 takes all three, plus Terminal-Bench 2.1 (88.3) and SWE Marathon (42.0) from Moonshot’s independently-consistent launch suite. If your product is an agent grinding through tools, terminals, and browsers, K3 is the stronger and cheaper model in this matchup.
Frontend development, full stop. The #1 Frontend Code Arena ranking is crowd-voted, blind, and independent — the most trustworthy signal in this entire comparison. K3 wins 6 of 7 frontend domains.
Price. 1.67× cheaper on every token tier, with Mooncake disaggregated inference pushing >90% cache-hit rates on coding workloads — meaning most real input bills land at the $0.30 tier.
Open weights. On July 27, 2026, K3’s full 2.8T weights release under a Modified MIT license. Self-hosting, deep fine-tuning, data residency, air-gapped deployment, zero vendor lock-in — Opus 5 will never offer any of these at any price. For a meaningful slice of companies, this single row of the table ends the comparison in K3’s favor.
You can test all of this yourself for free: K3 is live in the Kimi app with no credit card — an invite link gets you bonus membership credits — and our PC setup guide covers the desktop apps.
Context window and output limits
On paper it’s a tie: both models take a 1M-token context window (K3: 1,048,576; Opus 5: 1,000,000). The difference is economics and retention. Filling the window costs $3 with K3 ($0.30 on cache hits) vs $5 with Opus 5, and K3’s Kimi Delta Attention decodes long contexts up to 6.3× faster than its predecessor — with a 90.4 score on a full-1M-token eval using no context management at all. Opus 5 counters with its guaranteed 128K output and mature prompt caching. For whole-codebase analysis on a budget, K3 does the same job for 60% of the cost.
Speed and latency
K3 is measured at roughly 62 output tokens/sec with ~2s time-to-first-token. Opus 5’s speed is effort-dependent by design — fast at low, deliberate at max. Neither model is a sprinter at the settings where their benchmark scores were earned; both are thinking models built for long autonomous runs. If your UX needs instant chat responses, use Opus 5 at low effort — or K3 at low effort / k3-256k, both now available.
Open weights vs closed: the structural divide
Benchmarks will flip back and forth all year. This won’t:
- Kimi K3 becomes fully downloadable on July 27 — weights, technical report, Modified MIT license. Moonshot recommends supernode deployments (64+ accelerators), so this matters to companies and labs, not hobbyists. But for them it unlocks fine-tuning depth, data residency, and cost economics no API can match.
- Claude Opus 5 is closed forever. The trade: zero infrastructure decisions, best-in-class safety documentation, and distribution through every major cloud. Anthropic carries the operational burden; you carry the per-token bill and the lock-in.
Kimi K3 vs Claude Opus 5: which should you choose?
| Your situation | Pick |
|---|---|
| Agentic coding / terminal automation (DeepSWE, Terminal-Bench) | Kimi K3 |
| Frontend / UI development (#1 Arena, blind-voted) | Kimi K3 |
| Agentic web research (BrowseComp 91.2) | Kimi K3 |
| Business-process automation (raw pass rate) | Kimi K3 (30.8 vs 26.0) |
| Knowledge work & document analysis (GDPval-AA) | Claude Opus 5 (62.0) |
| Hardest reasoning problems (HLE) | Claude Opus 5 (64.7) |
| Open weights, self-hosting, fine-tuning | Kimi K3 (July 27) |
| 1M-context workloads on a budget | Kimi K3 |
| High-volume simple queries (effort ladder savings) | Claude Opus 5 |
| Enterprise compliance, AWS/GCP/Azure-native procurement | Claude Opus 5 |
| Long single-pass outputs (128K guaranteed) | Claude Opus 5 |
| Best default on pure price-performance | Kimi K3 |
The smartest move costs nothing: take your 10 hardest real tasks and run them through both. K3 is free to try right now, and Opus 5 is live on claude.ai. Vendor benchmarks will tell you who won their own test; your workload will tell you who wins yours.
FAQ
Is Kimi K3 better than Claude Opus 5? It depends on the task. On independent head-to-head benchmarks, K3 wins the agentic tests (AutomationBench 30.8 vs 26.0, BrowseComp 91.2 vs 90.8, DeepSWE 1.1 69.0 vs 68.8) and holds the #1 Frontend Code Arena ranking. Opus 5 wins knowledge work and hard reasoning (GDPval-AA 62.0 vs 55.6, Humanity’s Last Exam 64.7 vs 56.0). K3 is also ~1.7× cheaper on every token tier and ships open weights on July 27.
Is Kimi K3 cheaper than Claude Opus 5? Yes — about 1.7× cheaper on every tier: $3/$15 vs $5/$25 per million tokens (input/output), and $0.30 vs $0.50 for cached input. On a typical agentic workload mix that’s roughly $2.31 vs $3.71 per million blended tokens. Opus 5’s low-effort mode and 50% Batch API discount can narrow the gap for easy or offline workloads.
Kimi K3 vs Claude Opus 5 for coding — which wins? For agentic and frontend coding, K3: it edges DeepSWE 1.1 (69.0 vs 68.8), scores 88.3 on Terminal-Bench 2.1, and ranks #1 on the community-voted Frontend Code Arena. Anthropic claims state-of-the-art for Opus 5 on its own Frontier-Bench and CursorBench coding evals at a lower cost per task — those vendor numbers await independent confirmation.
What is Claude Opus 5’s “proactive” behavior? Anthropic describes Opus 5 as a “thoughtful and proactive” model — it takes initiative on ambiguous tasks, plans multi-step work, and self-verifies instead of waiting for instructions. It ships with thinking on by default and a five-level effort ladder (low through max) to trade intelligence for speed and cost per request.
Can I self-host Claude Opus 5 or Kimi K3? Only K3. Its full weights release July 27, 2026 under a Modified MIT license — the largest open-weight model ever at 2.8 trillion parameters. Opus 5 is closed and API-only via the Claude API, Amazon Bedrock, Google Cloud, and Microsoft Foundry.
Should I switch from Claude Opus 5 to Kimi K3? If your workload is agentic coding, terminal automation, frontend development, or agentic research — very likely yes, at about 60% of the token cost. If you need top knowledge-work scores, granular effort controls, 128K guaranteed outputs, or Anthropic’s compliance and cloud ecosystem, Opus 5 earns its premium. Benchmark both on your real tasks first.
The verdict
Kimi K3 vs Claude Opus 5 is the closest open-vs-closed fight of 2026. Opus 5 is Anthropic’s best efficiency play ever — genuinely near-Fable intelligence at half the price, with the strongest knowledge-work and reasoning scores in its class, an effort ladder for every budget, and enterprise distribution nobody else matches. It’s the best closed mainstream flagship you can buy.
But K3 wins the agentic categories that define where AI is actually heading — coding, automation, research — holds the only blind community-voted #1 in this comparison, costs 40% less per token, and in two days becomes the largest open-weight model ever released. For developers, startups, and anyone building agents at scale, K3 is the rational default. For regulated enterprises and knowledge-work-heavy teams, Opus 5 is the safest premium.
For the flagship-tier fight, see our Kimi K3 vs Claude Fable 5 breakdown; for the full four-way picture including GPT-5.6 Sol, read the frontier model comparison — or start from the complete Kimi K3 guide.
Try it yourself
Sign up for Kimi through our invite link and both of us get free bonus membership credits — up to a full year, at no cost to you.
Claim Free Credits →