On July 16, 2026, Moonshot AI launched Kimi K3 — a 2.8-trillion-parameter open-source model that rivals GPT-5.6 Sol and Claude Fable 5 on coding benchmarks. Here is the full pricing breakdown, benchmark results, and how to access kimi-k3 via API.
Kimi K3 — Moonshot AI's 2.8 trillion parameter open-weight model
Kimi K3 is Moonshot AI's 2.8-trillion-parameter model and the first announced open model in the 3T class. It uses a Stable LatentMoE architecture with 896 experts, activating 16 for each token, alongside Kimi Delta Attention and Attention Residuals. Moonshot announced that full weights would be released on July 27, 2026.
The model supports a native 1,048,576-token context window with flat pricing across the full window, rather than charging a higher rate for long prompts. It also accepts text, images, and video, making it relevant to coding agents that combine repository context with screenshots or other visual inputs.
Coding is the clearest strength in the launch results. K3 reached first place on LMArena's Frontend Code Arena in blind developer judging and led the reported SWE Marathon and BrowseComp comparisons. Those results make it a serious option for long-running coding sessions and large repositories, although benchmark methodology still matters.
Kimi K3 costs $3 per million uncached input tokens and $15 per million output tokens. Cached input is discounted to $0.30 per million tokens, and Moonshot reports a cache-hit rate above 90% for coding workloads on its official API. The model initially exposes max reasoning effort only, with lower and higher effort modes planned.
| Model | Input (cache miss) | Input (cache hit) | Output | Context |
|---|---|---|---|---|
| Kimi K3 (kimi-k3) | $3.00 | $0.30 | $15.00 | 1,048,576 |
Source: Moonshot AI official pricing as of July 16, 2026. Price is flat across the full context window. Max reasoning effort only (low/high modes coming). Official API reports >90% cache hit rate on coding workloads.
K3 is competitive with leading proprietary models, but it does not win every test. It leads the reported SWE Marathon and Program Bench results, nearly matches GPT-5.6 Sol on Terminal-Bench 2.1, and trails Sol on DeepSWE and Claude Fable 5 on FrontierSWE. BrowseComp and OmniDocBench are additional strengths in Moonshot's published comparison.
| Benchmark | Kimi K3 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|
| SWE Marathon | 42.0 | 39.0 | 35.0 |
| Program Bench | 77.8 | 77.6 | 76.8 |
| Terminal-Bench 2.1 | 88.3 | 88.8 | 84.6 |
| DeepSWE | 67.5 | 73.0 | 70.0 |
| FrontierSWE | 81.2 | 71.3 | 86.6 |
| BrowseComp | 91.2 | — | 88.0 |
| OmniDocBench | 91.1 | — | — |
Note: Benchmarks run in each vendor's own harness (KimiCode, Codex, Claude Code). K3 uses max reasoning. Fable 5 results may include Opus 4.8 fallback on some tasks.
K3's uncached input is 40% cheaper than GPT-5.6 Sol ($3 vs $5), while output is 50% cheaper ($15 vs $30). Compared with Claude Fable 5, K3's input and output rates are both 70% lower. Fable can still justify its premium when judgment and reliability matter more than raw token cost, but K3 is attractive for price-sensitive, long-context coding.
Prompt caching has the largest effect on sustained agent sessions. At a 90% input-cache hit rate, a workload using 10 million input tokens and 1 million output tokens costs about $20.70 on K3, compared with $39.50 on Sol and $69 on Fable 5 at the stated rates. Actual savings depend on whether the application preserves stable prompt prefixes and repository context between turns.
K3 is best suited to long-horizon coding, front-end generation, and agents that repeatedly work with a large, stable repository context. Its OpenAI-compatible API format also makes it easier to test through existing coding tools and model-routing platforms without adopting a separate SDK.
The launch has important limitations. Only max reasoning effort is available initially, and reported hallucination measurements rose from 39% on K2.6 to 51% on K3 despite higher answer accuracy. That tradeoff makes human review important for research or production decisions where a confident error is costly.
APItokendeal provides OpenAI-compatible access to supported Kimi, GPT, Claude, DeepSeek, and Gemini models through one key. Teams can evaluate K3 against other model families without maintaining a separate integration for each provider.
APItokendeal provides OpenAI-compatible API access to a wide range of models including Kimi K3, GPT, Claude, DeepSeek, Gemini, and more. One API key, one dashboard, prepaid token packs with no surprise bills.