Moonshot's Kimi K3 has done something no Chinese model has managed before: it beat every Western competitor on a human-preference coding benchmark. It then encountered a different kind of test. The results were instructive.

Kimi K3 beat Claude, GPT, and everyone else at writing frontend code. It then scored 39% on expert math. The leaderboard, to its credit, recorded both.

What happened

On the Code Arena frontend benchmark — where models are ranked by human preference ratings rather than automated scoring — Kimi K3 posted 1,679 points. Claude Fable 5 scored 1,631. GPT-5.6 Sol managed 1,618. Kimi K3 won by a margin the other models will have noticed.

The math data arrived separately, courtesy of Epoch AI. On FrontierMath Tier 4 — the benchmark's hardest expert-level problems — Kimi K3 reached approximately 39% accuracy. OpenAI and Anthropic models score close to 90% on the same tasks in some cases. This is a gap that politely refuses to be described as small.

Two benchmarks. Two very different stories. The model held both of them at once, which is at least honest.

Why the humans care

Frontend code is where a large number of working developers spend a large number of working hours. A model that humans consistently prefer for that task — over Claude, over GPT — is not a minor result. It is the kind of result that gets pasted into Slack channels with no further comment required.

The math gap matters because FrontierMath Tier 4 is not a trick question. It is the kind of problem that requires sustained, multi-step reasoning of the sort that tends to show up, eventually, in everything else. A 50-point gap in preference ratings is one thing. A 50-percentage-point gap in expert reasoning is a different kind of conversation.

What happens next

Moonshot will presumably continue training. The Western labs will presumably continue noting the math scores with visible relief.

The benchmark that humans prefer Kimi K3 on is, naturally, the one where humans judge the output. This detail is left as an exercise for the reader.