
Kimi K3 Put a Chinese Open Model at #1 on Frontend Code Arena for the First Time
Quick verdict
Moonshot released Kimi K3 and it is the first Chinese model that nobody in the field is dismissing. It scored 57 on the Artificial Analysis Intelligence Index, third overall behind Claude Fable 5 at 60 and GPT-5.6 at 59, and a step ahead of Claude Opus 4.8 at 56. On Frontend Code Arena it took the number-one spot at 1679, the first time a Chinese model has led that board, ahead of Fable 5 and GPT-5.6 Sol. It runs at roughly a third of Fable's price, and the weights are due out on July 27. The catch worth keeping in mind: it burns tokens, so the sticker price and the real cost per finished task are two different numbers.
What actually shipped
Kimi K3 is live now on kimi.com, the Kimi apps, Kimi Work, Kimi Code, and the API, with the default reasoning intensity set to max. The open weights are not out yet; Moonshot says the full model lands on July 27 along with a technical report. It is a very large mixture-of-experts model, with community estimates around 2.8T total parameters, and a 1M-token context window. That scale means the same thing it meant for Inkling and the other big open releases this year: this is not a model you run on a laptop, even after the weights drop.
The benchmark run is what got people talking. On the coding side, Artificial Analysis put K3 at 57 on its Coding Agent Index, matching GPT-5.6 Terra and GPT-5.5 and ahead of Opus 4.8, with 84% on Terminal-Bench v2, 64% on DeepSWE, and 23% on SWE-Atlas-QnA. DataCurve had it debut at #3 on DeepSWE, which it called the first open-weights model to post frontier-level results there. Guillermo Rauch's team reported it topping the Next.js evals, a first for an open model on that board. So the strength is concentrated in coding and frontend work, not spread evenly across every eval.
| Model | AA Intelligence Index | Frontend Code Arena |
|---|---|---|
| Claude Fable 5 | 60 | 1631 |
| GPT-5.6 Sol | 59 | 1618 |
| Kimi K3 | 57 | 1679 |
| Claude Opus 4.8 | 56 | — |
The architecture is doing real work
The piece that drew the most technical attention is Kimi Delta Attention. It works like a fast-weights memory: the model keeps a fixed-size learned state per request instead of paying the full attention cost across a long context. Moonshot claims up to 6x faster and cheaper throughput at 1M tokens, with pricing that stays flatter as the context grows instead of climbing the way it usually does. If that holds up outside their own numbers, it is one of the more important ideas here, because long-context cost is exactly where most agent workflows bleed money.
The bigger point people took from K3 is that the moat has moved. The story used to be about who had the most compute. K3 is a case that better MoE routing, quantization, and data curation can compress capability-per-FLOP enough to close the gap without matching Western spending directly. Whether it is genuinely at the frontier or a few months behind on the harder evals is still argued. What is settled is that you can no longer wave it off.
Why Anthropic should care more than OpenAI
Here is the number that matters past the headline rank. K3 is not as cheap in practice as the price tag suggests, because it uses roughly twice the tokens of GPT-5.6 Sol at high settings, which pushes its effective cost back up toward parity. Against Claude Opus 4.8, though, it runs at about half the tokens, so on a cost-per-task basis it undercuts Anthropic's flagship while scoring a point higher. That asymmetry is why the reaction framed K3 as a pricing problem for Anthropic specifically rather than for OpenAI.
Zoom out and the shape is familiar from earlier this year. A capable open-weight model shows up at a fraction of the closed price, usage starts routing toward it for the workloads where it is good enough, and the premium tiers have to justify their cost on the tasks that genuinely need them. If you are managing an AI bill and watching these releases land every few weeks, the practical move is the same as it has been: route the easy work to the cheap capable model and keep the flagship for what actually needs it. Our writeups on routing requests to the cheapest capable model and the best open-source AI models in 2026 both cover how to think about that split.
Video: a first look at Kimi K3
A walkthrough of what shipped, the benchmark claims, and how the open-weights plan fits.
FAQ
Can I download the Kimi K3 weights?
Not yet. Moonshot says the full weights land on July 27, 2026, alongside a technical report. Until then K3 is available through kimi.com, the Kimi apps, Kimi Code, Kimi Work, and the API. And even once the weights are out, community estimates put the model near 2.8T total parameters, so this is a hosted-inference model for almost everyone, not something you run at home.
Is Kimi K3 actually cheaper than Claude or GPT?
It depends on whether you count per token or per task. The per-token price is roughly a third of Fable's. But K3 uses about twice the tokens of GPT-5.6 Sol on hard problems, so against OpenAI the real cost lands closer to parity. Against Claude Opus 4.8 it uses about half the tokens, so there it is a genuine cost win.
What is Kimi Delta Attention?
It is a fast-weights memory mechanism that keeps a fixed-size learned state per request instead of paying full attention cost over a long context. Moonshot claims it delivers up to 6x faster and cheaper throughput at 1M tokens, with pricing that stays flatter as context grows. It is the main reason K3's long-context economics look different from most frontier models.
Sources
- Artificial Analysis - Kimi K3 at 57 on the Intelligence Index
- Artificial Analysis - K3 coding agent scores, Terminal-Bench and DeepSWE
- Arena - K3 leads Frontend Code Arena, first Chinese model to top it
- DataCurve - K3 debuts at #3 on DeepSWE
- @sdrzn - technical explainer on Kimi Delta Attention
- r/LocalLLaMA - Kimi K3 weights release date and availability
Further reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix