All models
Moonshot AI logo

Kimi-K3

By Moonshot AI · July 2026

Page last updated on

MoE1M contextMultimodal Kimi K3 License license

Open 2.8T-parameter MoE model from Moonshot AI with a 1M-token context window and native vision. Built for frontier coding, agentic work, and long-horizon reasoning.

Specifications

Total parameters
2.8T
Active parameters
104B
Architecture
MoE
Attention
Kimi Delta Attention (KDA) + Attention Residuals (AttnRes)
Context window
1M
Modality
Text, Image, Video
Precision
MXFP4 weights + MXFP8 activations
License
Kimi K3 License
Released
July 2026
Best for
Frontier coding, agentic search, and multimodal long-context reasoning

Good to know

  • The full Kimi K3 weights are now open. Moonshot AI released the model on Hugging Face under the Kimi K3 License, along with a technical report. K3 also runs through the Kimi app, Kimi Code, and the Kimi API as model kimi-k3, with max thinking effort by default. Day-0 serving support landed in vLLM and SGLang, and Moonshot contributed a Kimi Delta Attention kernel to the vLLM community.

  • Demand for Kimi K3 has run close to capacity since launch. Moonshot AI has temporarily paused new subscriptions to protect the experience for current members, and is adding compute to reopen new spots in batches. Existing subscribers are not affected. Membership is also being split into two plans: Kimi Membership for Kimi Web, App, and Work, and Kimi Code Membership for coding workflows.

Architecture

The headline change in Kimi K3 is the attention stack. The model is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two updates to how information moves across sequence length and model depth.

  • KDA is an efficient attention foundation that scales to the 1M-token context window. Moonshot contributed a KDA kernel to the vLLM community alongside the weights.
  • AttnRes selectively retrieves representations across depth, so deeper layers can reach back to earlier ones.

The Mixture-of-Experts layer is a Stable LatentMoE that activates 16 of 896 experts per token, with quantile balancing to stop experts from collapsing onto a few paths. Together these changes give roughly a 2.5x gain in scaling efficiency over Kimi K2. K3 uses quantization-aware training from the SFT stage onward, with MXFP4 weights and MXFP8 activations for broad hardware support. That keeps the KV cache small enough to serve a 1M-token context at a competitive token price, and the mixed-precision quantization holds the memory footprint down at 2.8T total parameters.

For more models like this one, browse the full open source LLM directory or read the open source LLM ecosystem statistics.

Benchmarks

Moonshot published a full benchmark set for K3 against the closed frontier. K3 trades the lead with Claude Fable 5 and GPT-5.6 Sol. It is strongest on agentic search and automation: it tops DeepSearchQA (95.0 F1), BrowseComp (91.2%), SWE Marathon (42.0%), and Automation Bench (30.8%). On coding it runs close to the best, at 88.3% on Terminal Bench 2.1 and 77.8% on Program Bench. On the vision side it leads OmniDocBench (91.1%).

Kimi K3 max vs frontier models

Scores published by Moonshot AI in the Kimi K3 announcement. Kimi K3 runs at max thinking effort. Elo-score rows (GDPval-AA v2, AA-Briefcase) are ratings, not percentages, so they render on their own axis. Higher is better on every benchmark. Scores marked n/a were not reported. Some competitor cells were evaluated by a third party.

Fable 5 max GPT-5.6 Sol max Opus 4.8 max GPT-5.5 xHigh GLM-5.2 max Kimi K3 max

Each model keeps the same color across every benchmark, stacked top to bottom in this order. Hover or focus a row to read all six scores.

DeepSWE
67.5 #3 of 6
Program Bench
77.8 best
Terminal Bench 2.1
88.3 #2 of 6
FrontierSWE
81.2 #2 of 6
SWE Marathon
42.0 best
PostTrain Bench
36.6 #2 of 6
MLS Bench
48.3 #2 of 6
View all scores as a table
Benchmark Fable 5 maxGPT-5.6 Sol maxOpus 4.8 maxGPT-5.5 xHighGLM-5.2 maxKimi K3 max
Coding
DeepSWE 70.0 73.0 59.0 67.0 46.2 67.5
Program Bench 76.8 77.6 71.9 70.8 63.7 77.8
Terminal Bench 2.1 84.6 88.8 84.6 83.4 82.7 88.3
FrontierSWE 86.6 71.3 66.7 64.9 67.3 81.2
SWE Marathon 35.0 39.0 40.0 14.0 13.0 42.0
PostTrain Bench 41.4 34.6 34.1 28.4 34.3 36.6
MLS Bench 49.9 46.2 42.8 35.5 40.4 48.3
Agentic
GDPval-AA v2 (Elo) 1,760 1,748 1,600 1,494 1,514 1,668
AA-Briefcase (Elo) 1,583 1,495 1,354 1,158 1,260 1,548
BrowseComp 88.0 90.4 84.3 84.4 n/a 91.2
DeepSearchQA (F1) 94.2 n/a 93.1 n/a n/a 95.0
Toolathlon-Verified 77.9 74.9 76.2 73.5 59.9 73.2
MCP Atlas 84.7 83.6 83.6 82.8 82.6 84.2
Automation Bench 29.1 29.7 27.2 22.7 12.9 30.8
Job Bench 57.4 46.5 48.4 38.3 43.4 52.9
APEX-Agents 43.3 39.9 39.4 38.5 35.6 41.0
Office QA Pro 69.9 63.2 63.9 60.9 41.4 63.3
SpreadsheetBench 2 34.7 32.4 31.6 29.1 28.1 34.8
Reasoning & Knowledge
GPQA-Diamond 92.6 94.1 91.0 93.5 91.2 93.5
HLE-Full 53.3 44.5 49.8 41.4 n/a 43.5
HLE-Full w/ tools 63.0 58.0 57.9 52.2 n/a 56.0
Vision
MMMU-Pro 81.2 83.0 78.9 81.2 n/a 81.6
CharXiv (RQ) 88.9 84.6 80.5 84.1 n/a 84.8
MathVision 94.8 95.8 86.7 92.2 n/a 94.3
ZeroBench_main (pass@5) 23.0 17.0 17.0 22.0 n/a 23.0
WorldVQA ForceAnswer 56.7 41.8 39.1 38.5 n/a 51.0
OmniDocBench 89.8 85.8 87.9 89.4 n/a 91.1
PerceptionBench 57.2 59.7 47.2 55.8 n/a 58.5

Source: the Kimi model card. Best score in each row is marked.

Third-party evaluations

Independent leaderboards started scoring K3 the week it was announced, ahead of the weight release. Across coding, agentic, and design arenas it lands at or near the top of the field, and it is the strongest open-weight entry in every one.

Arena Frontend Code Arena bar chart. Kimi-K3 ranks first at 1,679, ahead of Claude Fable 5 at 1,631, GPT-5.6 Sol (xHigh) at 1,618, and GLM-5.2 (Max) at 1,587.
Frontend Code Arena. On the Arena Frontend Code Arena leaderboard, Kimi-K3 ranks first at 1,679. It sits above Claude Fable 5 at 1,631 and GPT-5.6 Sol at 1,618. Arena · Jul 2026
Arena Code Arena Frontend bar chart of open-weight models. Kimi K3 (Max) ranks first at 1,682, ahead of GLM-5.2 (Max) at 1,587, GLM-5.1 at 1,518, Hy3 at 1,518, and Kimi-K2.6 at 1,510.
Frontend Code Arena, open-weight models. Filtered to open-weight models, Kimi K3 (Max) tops the Arena Code Arena Frontend board at 1,682. It leads GLM-5.2 (Max) at 1,587, GLM-5.1 and Hy3 at 1,518, and the earlier Kimi-K2.6 at 1,510. Arena · Jul 2026
Arena Agent Arena chart of net improvement versus baseline. Kimi K3 ranks fourth at +9.6%, behind Claude Fable 5 (High) at +13.2%, Claude Opus 4.8 (Thinking) at +10.0%, and GPT-5.6 Sol (xHigh) at +9.9%.
Agent Arena, top 25. Agent Arena scores models on live, long-horizon agentic sessions. Kimi K3 ranks fourth at +9.6% net improvement over baseline. That is a large jump from Kimi-K2.7 Code, which sits at -0.7%. Arena · Jul 2026
Artificial Analysis Coding Agent Index bar chart and cost-per-task scatter. Kimi K3 in the Kimi Code CLI scores 57, matching GPT-5.6 Terra and GPT-5.5, ahead of Opus 4.8 at 55, GLM-5.2 at 40, and DeepSeek V4 Pro at 28.
Artificial Analysis Coding Agent Index. On the Artificial Analysis Coding Agent Index, K3 in the Kimi Code CLI scores 57, joint fourth overall. It matches GPT-5.6 Terra and GPT-5.5, and beats Opus 4.8 at 55. Among open-weight entries it leads by a wide margin, well ahead of GLM-5.2 at 40 and DeepSeek V4 Pro at 28. The scatter plot puts K3 in the most attractive cost-per-task quadrant at about $3.18 per task. Artificial Analysis · Jul 2026
Vals AI Vals Index. Kimi K3 scores 74.70% and ranks 2 of 38, above Claude Opus 4.8 at 70.36% and GPT-5.6 Sol at 73.12%, just behind Claude Fable 5 at 75.14%.
Vals Index. In the Vals Index, K3 scores 74.70% and ranks second of 38 models. It beats GPT-5.6 Sol at 73.12% and Claude Opus 4.8 at 70.36%. Only Claude Fable 5 is ahead at 75.14%. It is a big step up from Kimi K2.6 at 55.17%. Vals AI · Jul 16, 2026
Vals AI Vibe Code Bench v1.1 table. Kimi K3 ranks second at 84.96% accuracy on the OpenHands harness, behind Claude Fable 5 at 90.35% and ahead of Claude Opus 4.8 at 82.72%.
Vibe Code Bench v1.1. The Vals AI Vibe Code Bench asks models to build web apps from scratch. K3 ranks second at 84.96%, behind Claude Fable 5 at 90.35% and ahead of Claude Opus 4.8 at 82.72%. Vals AI · Jul 18, 2026
DesignArena 3D Design Elo chart. Kimi K3 ranks first at 1,450, ahead of Claude Fable 5 at 1,368, GLM-5.2 at 1,363, and GPT-5.6 Sol at 1,359.
DesignArena, 3D Design. On the DesignArena 3D Design leaderboard, K3 ranks first at 1,450 Elo. It sits above Claude Fable 5 at 1,368 and GLM-5.2 at 1,363. DesignArena · Jul 2026

API pricing

K3 is served through the official Kimi API as model kimi-k3, with max thinking effort by default. Output is $15.00 per million tokens. Cache-miss input is $3.00 per million tokens, and a cache hit drops input to $0.30 per million tokens. Moonshot credits the Mooncake disaggregated inference stack for the low cached price.

Model Input · cache hit Input · cache miss Output
Kimi-K3 $0.30 $3.00 $15.00

Per 1M tokens, as of July 22, 2026. Official pricing.

For a 2.8T-parameter model, that output price is competitive. It undercuts Claude Opus 4.8 at $25 per million output tokens. On the Artificial Analysis coding runs, K3 averaged about $3.18 per task, well inside the most attractive cost-per-task quadrant and cheaper than Opus 4.8 and GPT-5.6 Sol.

Sources: Kimi K3 announcement · Claude pricing

What people are saying

“Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5).”

Elon Musk

@elonmusk · xAI

X Jul 2026

“Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete his Ph.D. in just four years, but he also made truly fundamental contributions to ML during his time at CMU.”

Russ Salakhutdinov

@rsalakhu · Carnegie Mellon University

X Jul 2026

“It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026.”

Dean W. Ball

@deanwball · OpenAI

X Jul 2026

“Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5).”

Elon Musk

@elonmusk · xAI

X Jul 2026

“Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete his Ph.D. in just four years, but he also made truly fundamental contributions to ML during his time at CMU.”

Russ Salakhutdinov

@rsalakhu · Carnegie Mellon University

X Jul 2026

“It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026.”

Dean W. Ball

@deanwball · OpenAI

X Jul 2026

Anecdotes

A game clone for the price of a coffee

Developer Chris (@ChrisGPT) had Kimi K3 build a CS:GO x Portal clone in a single run. It used around 600,000 tokens and cost about $3.24 in API usage. He noted the same job would run roughly $10.80 on Claude Fable 5 and about $6 on GPT-5.6 Sol, and called it a sign that low-cost indie game development is getting close.

Chris on X

The Moonshot CEO helped write the original GLM paper

A twist in the open-model rivalry made the rounds after launch. Zhilin Yang, the co-founder and CEO of Moonshot AI, was a co-author on the 2022 paper 'GLM: General Language Model Pretraining with Autoregressive Blank Infilling', published at ACL 2022. GLM is the lineage that later produced GLM-5.2, one of the models K3 is benchmarked against.

Zhilin Yang publications

Run it yourself

vllm serve moonshotai/Kimi-K3
python3 -m sglang.launch_server --model-path moonshotai/Kimi-K3

Frequently asked questions

Is Kimi K3 open source?

Yes. Moonshot AI released the full Kimi K3 weights on Hugging Face under the Kimi K3 License, along with a technical report. You can download and self-host the model, and day-0 serving support is available in vLLM and SGLang.

How big is Kimi K3?

Kimi K3 has 2.8 trillion total parameters, which Moonshot calls the first open 3T-class model. It is a Mixture-of-Experts model that activates 104 billion parameters per token, 16 of 896 experts, so only a small fraction of the parameters run on any given token.

What can Kimi K3 do?

Kimi K3 understands text, images, and video natively within one model, and supports a 1-million-token context window. It is built for frontier coding, agentic search and automation, and long-horizon reasoning. It ranks first on the Arena Frontend Code Arena and fourth on Agent Arena.

How much does the Kimi K3 API cost?

Through the official Kimi API, Kimi K3 costs $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens for output.

How does Kimi K3 compare to Claude Fable 5 and GPT-5.6 Sol?

Kimi K3 trades the lead with them. It leads on agentic search tasks like DeepSearchQA and BrowseComp and on SWE Marathon, while Fable 5 and GPT-5.6 Sol lead on several other coding and vision tests. On the Vals Index, K3 scores 74.70%, just behind Fable 5 at 75.14% and ahead of GPT-5.6 Sol at 73.12%.

Keep exploring