Kimi-K3
By Moonshot AI · July 2026
Page last updated on
Open 2.8T-parameter MoE model from Moonshot AI with a 1M-token context window and native vision. Built for frontier coding, agentic work, and long-horizon reasoning.
Specifications
- Total parameters
- 2.8T
- Active parameters
- 104B
- Architecture
- MoE
- Attention
- Kimi Delta Attention (KDA) + Attention Residuals (AttnRes)
- Context window
- 1M
- Modality
- Text, Image, Video
- Precision
- MXFP4 weights + MXFP8 activations
- License
- Kimi K3 License
- Released
- July 2026
- Best for
- Frontier coding, agentic search, and multimodal long-context reasoning
Good to know
-
The full Kimi K3 weights are now open. Moonshot AI released the model on Hugging Face under the Kimi K3 License, along with a technical report. K3 also runs through the Kimi app, Kimi Code, and the Kimi API as model kimi-k3, with max thinking effort by default. Day-0 serving support landed in vLLM and SGLang, and Moonshot contributed a Kimi Delta Attention kernel to the vLLM community.
-
Demand for Kimi K3 has run close to capacity since launch. Moonshot AI has temporarily paused new subscriptions to protect the experience for current members, and is adding compute to reopen new spots in batches. Existing subscribers are not affected. Membership is also being split into two plans: Kimi Membership for Kimi Web, App, and Work, and Kimi Code Membership for coding workflows.
Architecture
The headline change in Kimi K3 is the attention stack. The model is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), two updates to how information moves across sequence length and model depth.
- KDA is an efficient attention foundation that scales to the 1M-token context window. Moonshot contributed a KDA kernel to the vLLM community alongside the weights.
- AttnRes selectively retrieves representations across depth, so deeper layers can reach back to earlier ones.
The Mixture-of-Experts layer is a Stable LatentMoE that activates 16 of 896 experts per token, with quantile balancing to stop experts from collapsing onto a few paths. Together these changes give roughly a 2.5x gain in scaling efficiency over Kimi K2. K3 uses quantization-aware training from the SFT stage onward, with MXFP4 weights and MXFP8 activations for broad hardware support. That keeps the KV cache small enough to serve a 1M-token context at a competitive token price, and the mixed-precision quantization holds the memory footprint down at 2.8T total parameters.
For more models like this one, browse the full open source LLM directory or read the open source LLM ecosystem statistics.
Benchmarks
Moonshot published a full benchmark set for K3 against the closed frontier. K3 trades the lead with Claude Fable 5 and GPT-5.6 Sol. It is strongest on agentic search and automation: it tops DeepSearchQA (95.0 F1), BrowseComp (91.2%), SWE Marathon (42.0%), and Automation Bench (30.8%). On coding it runs close to the best, at 88.3% on Terminal Bench 2.1 and 77.8% on Program Bench. On the vision side it leads OmniDocBench (91.1%).
Kimi K3 max vs frontier models
Scores published by Moonshot AI in the Kimi K3 announcement. Kimi K3 runs at max thinking effort. Elo-score rows (GDPval-AA v2, AA-Briefcase) are ratings, not percentages, so they render on their own axis. Higher is better on every benchmark. Scores marked n/a were not reported. Some competitor cells were evaluated by a third party.
Each model keeps the same color across every benchmark, stacked top to bottom in this order. Hover or focus a row to read all six scores.
Separate axis. Elo is not a percentage, so this row is plotted as points on a fitted scale starting at 1,400.
Separate axis. Elo is not a percentage, so this row is plotted as points on a fitted scale starting at 1,000.
- Fable 5 max
- GPT-5.6 Sol max
- Opus 4.8 max
- GPT-5.5 xHigh
- GLM-5.2 max
- Kimi K3 max
View all scores as a table
| Benchmark | Fable 5 max | GPT-5.6 Sol max | Opus 4.8 max | GPT-5.5 xHigh | GLM-5.2 max | Kimi K3 max |
|---|---|---|---|---|---|---|
| Coding | ||||||
| DeepSWE | 70.0 | 73.0 | 59.0 | 67.0 | 46.2 | 67.5 |
| Program Bench | 76.8 | 77.6 | 71.9 | 70.8 | 63.7 | 77.8 |
| Terminal Bench 2.1 | 84.6 | 88.8 | 84.6 | 83.4 | 82.7 | 88.3 |
| FrontierSWE | 86.6 | 71.3 | 66.7 | 64.9 | 67.3 | 81.2 |
| SWE Marathon | 35.0 | 39.0 | 40.0 | 14.0 | 13.0 | 42.0 |
| PostTrain Bench | 41.4 | 34.6 | 34.1 | 28.4 | 34.3 | 36.6 |
| MLS Bench | 49.9 | 46.2 | 42.8 | 35.5 | 40.4 | 48.3 |
| Agentic | ||||||
| GDPval-AA v2 (Elo) | 1,760 | 1,748 | 1,600 | 1,494 | 1,514 | 1,668 |
| AA-Briefcase (Elo) | 1,583 | 1,495 | 1,354 | 1,158 | 1,260 | 1,548 |
| BrowseComp | 88.0 | 90.4 | 84.3 | 84.4 | n/a | 91.2 |
| DeepSearchQA (F1) | 94.2 | n/a | 93.1 | n/a | n/a | 95.0 |
| Toolathlon-Verified | 77.9 | 74.9 | 76.2 | 73.5 | 59.9 | 73.2 |
| MCP Atlas | 84.7 | 83.6 | 83.6 | 82.8 | 82.6 | 84.2 |
| Automation Bench | 29.1 | 29.7 | 27.2 | 22.7 | 12.9 | 30.8 |
| Job Bench | 57.4 | 46.5 | 48.4 | 38.3 | 43.4 | 52.9 |
| APEX-Agents | 43.3 | 39.9 | 39.4 | 38.5 | 35.6 | 41.0 |
| Office QA Pro | 69.9 | 63.2 | 63.9 | 60.9 | 41.4 | 63.3 |
| SpreadsheetBench 2 | 34.7 | 32.4 | 31.6 | 29.1 | 28.1 | 34.8 |
| Reasoning & Knowledge | ||||||
| GPQA-Diamond | 92.6 | 94.1 | 91.0 | 93.5 | 91.2 | 93.5 |
| HLE-Full | 53.3 | 44.5 | 49.8 | 41.4 | n/a | 43.5 |
| HLE-Full w/ tools | 63.0 | 58.0 | 57.9 | 52.2 | n/a | 56.0 |
| Vision | ||||||
| MMMU-Pro | 81.2 | 83.0 | 78.9 | 81.2 | n/a | 81.6 |
| CharXiv (RQ) | 88.9 | 84.6 | 80.5 | 84.1 | n/a | 84.8 |
| MathVision | 94.8 | 95.8 | 86.7 | 92.2 | n/a | 94.3 |
| ZeroBench_main (pass@5) | 23.0 | 17.0 | 17.0 | 22.0 | n/a | 23.0 |
| WorldVQA ForceAnswer | 56.7 | 41.8 | 39.1 | 38.5 | n/a | 51.0 |
| OmniDocBench | 89.8 | 85.8 | 87.9 | 89.4 | n/a | 91.1 |
| PerceptionBench | 57.2 | 59.7 | 47.2 | 55.8 | n/a | 58.5 |
Source: the Kimi model card. Best score in each row is marked.
Third-party evaluations
Independent leaderboards started scoring K3 the week it was announced, ahead of the weight release. Across coding, agentic, and design arenas it lands at or near the top of the field, and it is the strongest open-weight entry in every one.
API pricing
K3 is served through the official Kimi API as model kimi-k3, with max thinking effort by default. Output is $15.00 per million tokens. Cache-miss input is $3.00 per million tokens, and a cache hit drops input to $0.30 per million tokens. Moonshot credits the Mooncake disaggregated inference stack for the low cached price.
| Model | Input · cache hit | Input · cache miss | Output |
|---|---|---|---|
| Kimi-K3 | $0.30 | $3.00 | $15.00 |
Per 1M tokens, as of July 22, 2026. Official pricing.
For a 2.8T-parameter model, that output price is competitive. It undercuts Claude Opus 4.8 at $25 per million output tokens. On the Artificial Analysis coding runs, K3 averaged about $3.18 per task, well inside the most attractive cost-per-task quadrant and cheaper than Opus 4.8 and GPT-5.6 Sol.
Sources: Kimi K3 announcement · Claude pricing
What people are saying
“Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5).”
Elon Musk
@elonmusk · xAI
X Jul 2026
“Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete his Ph.D. in just four years, but he also made truly fundamental contributions to ML during his time at CMU.”
Russ Salakhutdinov
@rsalakhu · Carnegie Mellon University
X Jul 2026
“It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026.”
Dean W. Ball
@deanwball · OpenAI
X Jul 2026
“Our 2T model, which is better than our 1.5T in every way, will finish initial training next week. It might be able to exceed Kimi, but with speed and token efficiency close to our 1.5T (aka Grok 4.5).”
Elon Musk
@elonmusk · xAI
X Jul 2026
“Congratulations to Zhilin Yang, founder and CEO of @Kimi_Moonshot, on the latest Kimi release. What a huge win for the open-source community! It feels like just yesterday Zhilin was graduating from my lab at CMU, jointly co-advised with William Cohen. Not only did he complete his Ph.D. in just four years, but he also made truly fundamental contributions to ML during his time at CMU.”
Russ Salakhutdinov
@rsalakhu · Carnegie Mellon University
X Jul 2026
“It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026.”
Dean W. Ball
@deanwball · OpenAI
X Jul 2026
Anecdotes
A game clone for the price of a coffee
Developer Chris (@ChrisGPT) had Kimi K3 build a CS:GO x Portal clone in a single run. It used around 600,000 tokens and cost about $3.24 in API usage. He noted the same job would run roughly $10.80 on Claude Fable 5 and about $6 on GPT-5.6 Sol, and called it a sign that low-cost indie game development is getting close.
The Moonshot CEO helped write the original GLM paper
A twist in the open-model rivalry made the rounds after launch. Zhilin Yang, the co-founder and CEO of Moonshot AI, was a co-author on the 2022 paper 'GLM: General Language Model Pretraining with Autoregressive Blank Infilling', published at ACL 2022. GLM is the lineage that later produced GLM-5.2, one of the models K3 is benchmarked against.
Run it yourself
vllm serve moonshotai/Kimi-K3 python3 -m sglang.launch_server --model-path moonshotai/Kimi-K3 Frequently asked questions
Is Kimi K3 open source?
Yes. Moonshot AI released the full Kimi K3 weights on Hugging Face under the Kimi K3 License, along with a technical report. You can download and self-host the model, and day-0 serving support is available in vLLM and SGLang.
How big is Kimi K3?
Kimi K3 has 2.8 trillion total parameters, which Moonshot calls the first open 3T-class model. It is a Mixture-of-Experts model that activates 104 billion parameters per token, 16 of 896 experts, so only a small fraction of the parameters run on any given token.
What can Kimi K3 do?
Kimi K3 understands text, images, and video natively within one model, and supports a 1-million-token context window. It is built for frontier coding, agentic search and automation, and long-horizon reasoning. It ranks first on the Arena Frontend Code Arena and fourth on Agent Arena.
How much does the Kimi K3 API cost?
Through the official Kimi API, Kimi K3 costs $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens for output.
How does Kimi K3 compare to Claude Fable 5 and GPT-5.6 Sol?
Kimi K3 trades the lead with them. It leads on agentic search tasks like DeepSearchQA and BrowseComp and on SWE Marathon, while Fable 5 and GPT-5.6 Sol lead on several other coding and vision tests. On the Vals Index, K3 scores 74.70%, just behind Fable 5 at 75.14% and ahead of GPT-5.6 Sol at 73.12%.