Moonshot AI has launched Kimi K3, which they claim is the world’s largest open AI model. The announcement is another major development in the global AI race as the Chinese startup is aiming to narrow the gap with leading U.S. AI companies. Unlike closed AI models, Kimi K3 is released as an open weight model, which gives developers more freedom to use, customize and deploy it for different applications.
While leaks suggested 2.5 trillion parameters, the mode is now confirmed to be built using a 2.8 trillion parameter Mixture-of-Experts (MoE) architecture and supports a 1 million token context window. These numbers make it one of the largest and most capable open AI models available today. The launch is significant because it combines large scale performance with an open approach, offering developers alongside proprietary models.
What’s New and How to Get Started with Kimi K3
Kimi K3 brings several upgrades for developers building AI driven tools. Here are some highlights:
- Built with a 2.8 trillion parameter Mixture-of-Experts (MoE) architecture.
- Supports a 1 million token context window for handling long documents, conversations and code.
- Released as an open weight model, allowing developers to self host and customize it.
- Designed for coding, reasoning and AI agent tasks.
- Uses an OpenAI-compatible API, making it easier to migrate existing applications.
Getting started with Kimi K3 is simple. Developers can create an API key through the Kimi platform and connect to the official API endpoint. Since the API follows an OpenAI-compatible format, many existing applications require only minor configuration changes before they can start using Kimi K3.
Kimi K3 currently sits at number one on the Frontend Code Arena leaderboard:
Big news: Kimi-K3 by @Kimi_Moonshot is now #1 in the Frontend Code Arena with 1679 pts, surpassing Claude Fable 5.
— Arena.ai (@arena) July 16, 2026
This is a 17-place jump from Kimi-k2.6 (#18 -> #1).
In Frontend, Kimi-K3 ranked #1 in 6 of 7 domains: Brand & Marketing, Reference-Based Design, Data & Analytics,… https://t.co/YDN3BufGkC pic.twitter.com/Oa6teaQnWp
How it Compares to GPT-5.6 Sol, Mythos and Fable 5
With Kimi K3, Moonshot is targeting the same category of advanced AI models as GPT-5.6 Sol, Claude Opus 4.8, GPT-5.5, GLM-5.2 and Claude Fable 5. The company is positioning K3 as a frontier level model capable of competing in coding, reasoning and AI agent tasks.
Benchmark Comparison
| BENCHMARKS | KIMI K3 | CLAUDE FABLE 5 | GPT-5.6 SOL | CLAUDE OPUS 4.8 | GPT-5.5 | GLM-5.2 |
| DeepSWE | 67.5 | 70.0 | 73.0 | 59.0 | 67.0 | 46.2 |
|---|---|---|---|---|---|---|
| Program Bench | 77.8 | 76.8 | 77.6 | 71.9 | 70.8 | 63.7 |
| Terminal Bench 2.1 | 88.3 | 84.6 | 88.8 | 84.6 | 83.4 | 82.7 |
| FrontierSWE | 81.2 | 86.6 | 71.3 | 66.7 | 64.9 | 67.3 |
| SWE Marathon | 42.0 | 35.0 | 39.0 | 40.0 | 14.0 | 13.0 |
| PostTrain Bench | 36.6 | 41.4 | 34.6 | 34.1 | 28.4 | 34.3 |
| MLS Bench | 48.3 | 49.9 | 46.2 | 42.8 | 35.5 | 40.4 |
| Kimi Code Bench 2.0 (Internal) | 72.9 | 76.9 | 64.8 | 71.7 | 69.0 | 64.2 |
| GDPval-AA v2 (Elo-score) | 1668.0 | 1760.0 | 1748.0 | 1600.0 | 1494.0 | 1514.0 |
| BrowseComp | 91.2 | 88.0 | 90.4 | 84.3 | 84.4 | — |
| DeepSearchQA (f1-score) | 95.0 | 94.2 | — | 93.1 | — | — |
| Toolathlon-Verified | 73.2 | 77.9 | 74.9 | 76.2 | 73.5 | 59.9 |
| MCP Atlas | 84.2 | 84.7 | 83.6 | 83.6 | 82.8 | 82.6 |
| Automation Bench | 30.8 | 29.1 | 29.7 | 27.2 | 22.7 | 12.9 |
| Job Bench | 52.9 | 57.4 | 46.5 | 48.4 | 38.3 | 43.4 |
| AA-Briefcase (Elo-score) | 1548.0 | 1583.0 | 1495.0 | 1354.0 | 1158.0 | 1260.0 |
| APEX-Agents | 37.6 | 43.3 | 39.9 | 39.4 | 38.5 | 35.6 |
| Office QA Pro | 63.3 | 69.9* | 63.2* | 63.9* | 60.9* | 41.4 |
| SpreadsheetBench 2 | 34.8 | 34.7* | 32.4* | 31.6* | 29.1* | 28.1 |
| DECK-Bench (Internal) | 73.5 | 73.0 | 74.7 | 66.9 | 68.2 | 68.6 |
| GPQA-Diamond | 93.5 | 92.6 | 94.1 | 91.0 | 93.5 | 91.2 |
| HLE-Full | 43.5 | 53.3 | 44.5 | 49.8* | 41.4* | — |
| HLE-Full w/ tools | 56.0 | 63.0 | 58.0 | 57.9* | 52.2* | — |
| MMMU-Pro | 81.6 | 81.2 | 83.0 | 78.9 | 81.2 | — |
| MMMU-Pro w/ python | 83.4 | 86.5 | 84.6 | 82.7 | 83.2 | — |
| CharXiv (RQ) | 84.8 | 88.9 | 84.6 | 80.5 | 84.1 | — |
| CharXiv (RQ) w/ python | 91.3 | 93.5 | 89.1 | 89.9 | 89.0 | — |
| MathVision | 94.3 | 94.8 | 95.8 | 86.7 | 92.2 | — |
| MathVision w/ python | 97.8 | 98.6 | 97.8 | 97.1 | 96.8 | — |
| BabyVision w/ python | 85.7 | 90.5 | 88.9 | 81.2 | 83.6 | — |
| ZeroBench_main (pass@5) | 23.0 | 23.0 | 17.0 | 17.0 | 22.0 | — |
| ZeroBench_main w/ python (pass@5) | 41.0 | 46.0 | 35.0 | 34.0 | 41.0 | — |
| WorldVQA ForceAnswer | 51.0 | 56.7 | 41.8 | 39.1 | 38.5 | — |
| OmniDocBench | 91.1 | 89.8 | 85.8 | 87.9 | 89.4 | — |
| PerceptionBench | 58.5 | 57.2 | 59.7 | 47.2 | 55.8 | — |
Table Credits: Kimi AI
Here’s the link to another comparison table for better understanding: OpenRouter (Claude Fable 5 vs GPT-5.6 Sol vs Kimi K3 vs Claude Opus 4.8 vs GPT-5.5)
Based on the benchmark results above, each model appears to have its own strengths. Claude Fable 5 appears to be built for advanced reasoning, coding and research. GPT-5.6 Sol appears to perform well across writing, coding and business tasks. Claude Opus 4.8 seems designed for complex reasoning and detailed work. GPT-5.5 looks like a dependable option for everyday AI tasks, while GLM 5.2 appears to be another strong model that focuses on reasoning and multilingual support.
Kimi K3, though, looks promising for combining strong performance with lower API costs, which could make it a practical choice for developers who want good results without spending as much on inference. However, these observations are based on the available benchmark data, and actual performance may vary depending on the task and how the model is used.

Benchmarks are useful for measuring model performance, but they are only one part of the picture. Cost also plays a major role, especially for businesses. According to Moonshot’s official pricing page, Kimi K3 costs $3.00 per million input tokens (cache miss), $0.30 per million cached input tokens (cache hit), and $15.00 per million output tokens. These prices make it one of the most affordable frontier AI models currently available.
Also read: What We Know About Moonshot AI’s Unannounced Kimi K3
The launch of Kimi K3 marks an important moment for both Moonshot AI and the open AI community. The model shows how open AI systems are becoming more capable of competing with leading proprietary models. It also gives developers another high performance option that is easier to customize and deploy. OpenAI’s GPT and Claude models continue to lead many major AI workloads, but Kimi K3 brings together strong capabilities, competitive pricing and deployment flexibility in a single package. As businesses and developers continue to evaluate AI models based on performance, efficiency and cost, Kimi K3 is likely to become an important option for those looking to build advanced AI applications without relying solely on closed models.









