← All Posts

Kimi K3 Benchmarks: How Moonshot's New Model Actually Scores

By OpenClaw Launch Team

A Quiet Overnight Release

Kimi K3 showed up on kimi.com overnight on July 16, 2026, with no keynote and no technical report. Moonshot AI shipped two variants at once — K3 Max for chat and agentic work, and K3 Swarm Max for parallel processing — and let the model speak for itself. That is unusual for a release this size: at 2.8 trillion total parameters (a mixture-of-experts design with 896 experts and 16 active per token), it is being called the largest open-weight-class model announced to date. Moonshot has not disclosed the active-parameter count, and is reportedly raising at a roughly $31.5B valuation.

Artificial Analysis: Where K3 Actually Ranks

Independent numbers from Artificial Analysis put Kimi K3 at roughly 57 on its Intelligence Index — good for third place, behind Claude Fable 5 (~60) and GPT-5.6 Sol (~59), and ahead of GPT-5.5, Opus 4.8, and GLM-5.2 (51) and DeepSeek V4 Pro (44). On Artificial Analysis's long-horizon agentic eval, K3 posts an Elo of 1547, trailing only Fable 5. It also takes the top spot on Frontend Code Arena, ahead of Fable 5 — the first time a Kimi release has led that board outright.

Cost is where K3 stands out on Artificial Analysis's own accounting: an average of $0.94 per completed task, versus $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8. Moonshot also says K3 uses 21% fewer output tokens than K2.6 on comparable tasks, which is part of why the per-task cost is low despite a similar per-token price to competitors.

Moonshot's Own Harness

Moonshot's self-reported numbers tell a similar story but with different tools: Terminal-Bench 2.1 at 88.3 (versus 84.6 for Fable 5), SWE Marathon at 42.0 (versus 35.0 for Fable 5), plus 67.5 on DeepSWE and 81.2 on FrontierSWE. Moonshot's own summary of the release is blunt about where it sits in the pack: the company describes K3 as “mostly beating Claude Opus 4.8 and GPT-5.5, while losing out to Claude Fable 5 and GPT-5.6 Sol.” No MMLU number has been published, and SWE-bench Verified figures circulating for K3 disagree across sources depending on harness — we are not printing one until Moonshot or a neutral evaluator settles on it.

Why the Numbers Don't Line Up Cleanly

Two organizations testing the same model can land 10 to 26 points apart on SWE-bench-style benchmarks purely from harness differences — sandbox setup, retry policy, timeout length, and scoring rubric all move the number. That is true industry-wide, not a Kimi-specific problem, but it is a useful reason to read any single benchmark chart skeptically and to weight the Artificial Analysis Intelligence Index (a composite across many evals) over any one leaderboard screenshot. For a first read on where K3 sits: strong on coding and terminal/agentic tasks, competitive but not leading on aggregate intelligence, and meaningfully cheaper per task than the two models ranked above it.

Trying K3 Yourself

Kimi K3 is live now on the Moonshot API and on OpenRouter (watch for 429s — upstream capacity is limited right now). For what it costs to actually run and how kimi.com's subscriptions differ from API access, see our Kimi K3 pricing breakdown. Both OpenClaw and Hermes Agent can run K3 as the model behind an always-on chat agent on Telegram, Discord, WhatsApp, WeChat, or the web — bring your own Moonshot or OpenRouter key on OpenClaw Launch and the model is a dropdown, not a redeploy.

Build with OpenClaw

Deploy your own AI agent in under 30 seconds — no servers, no CLI.

Deploy Now