UK AISI + CAISI · 2026-07-23 · major
UK AISI + CAISI evaluate Kimi K3 — 32% on ExploitBench vs 76% for US frontier models
UK AISI and NIST's CAISI evaluated Kimi K3 on cyber tasks: 32% on ExploitBench vs 76% for US frontier models, 0 of 41 arbitrary-code-execution outcomes, and step 17 of a 32-step network attack vs 28.5 for the frontier.
.png)
UK AISI and NIST's CAISI put Kimi K3 through joint cyber evals and found it trails the US frontier by a wide margin — but still leads open-weight rivals.
Key specs
| Exploit bench (kimi k3) | 32% |
|---|---|
| Exploit bench (top us models) | 76% |
| Kimi k3 ace (of 41 tasks) | 0 |
| Network attack steps (kimi k3) | 17 of 32 |
| Network attack steps (frontier avg) | 28.5 |
Quick facts
| Evaluators | UK AISI + US CAISI (NIST) |
|---|---|
| Model tested | Kimi K3 (Moonshot AI, released 16 Jul 2026) |
| Cyber-exploit benchmark | ExploitBench (CMU) — 41 Chrome V8 vulnerabilities post-2023 |
| Kimi K3 ExploitBench score | 32% (vs 76% for top US models) |
| Arbitrary code execution | 0 of 41 samples (US frontier avg: 20 of 41) |
| Simulated network attack | step 17 of 32 avg (US frontier avg: 28.5) |
| vs GLM-5.2 | Kimi K3 32% beats GLM-5.2 24% on ExploitBench |
| Safety alignment | K3 did not refuse offensive-cyber tasks |
Benchmarks
| Kimi K3 (Moonshot) | 32% | |
|---|---|---|
| US frontier models (avg) | 76% | |
| GLM-5.2 (Z.ai) | 24% |
What is it?
The UK Artificial Intelligence Security Institute and the US Center for AI Standards and Innovation, part of NIST, published a joint preliminary assessment of Moonshot AI's Kimi K3 on offensive cyber tasks. It is the first cross-Atlantic government evaluation of a Chinese frontier open-weight model on this axis.
How does it work?
The two agencies ran Kimi K3 through Carnegie Mellon's ExploitBench, a set of 41 exploit-development tasks against Chrome V8 vulnerabilities disclosed after 2023, and through a 32-step simulated network-attack range. Each task was scored against the current US frontier closed models and the strongest open-weight model as of June, GLM-5.2, using an Item Response Theory methodology the labs have published in prior reports.
Why does it matter?
The result gives policymakers concrete numbers instead of vibes: Kimi K3 is not at frontier cyber capability, but it is the new open-weight ceiling and its safeguards do not refuse offensive-cyber help. The evaluation lands four days before Moonshot's scheduled 27 July open-weight drop, and while the White House weighs Entity-List designations against Moonshot — giving a shared UK-US technical baseline for those talks.
Who is it for?
Security teams, policy watchers, and open-weight model users trying to size the real cyber-abuse risk of a new frontier Chinese model.
Frequently asked questions
- What did UK AISI and CAISI actually measure with Kimi K3?
- UK AISI and CAISI ran Kimi K3 through ExploitBench, a Carnegie Mellon University benchmark that asks a model to write exploits for 41 Chrome V8 engine vulnerabilities disclosed after 2023, plus a 32-step simulated network-attack range. They compared its results against the current US frontier closed models and against the strongest open-weight alternatives.
- How does Kimi K3 compare to US frontier models on cyber tasks?
- Kimi K3 sits well behind. On ExploitBench it scored 32% to the leading US models' 76%. On arbitrary code execution it landed 0 of 41 tasks compared with a US-frontier average of 20 of 41. On the network-attack range it made step 17 of 32 on average, versus 28.5 steps for the strongest models — and completed the full attack in only 1 of 10 attempts.
- How does Kimi K3 compare to other open-weight models?
- Kimi K3 is now the strongest open-weight model on this cyber suite. Its 32% ExploitBench score beats Z.ai's GLM-5.2 at 24%, which UK AISI had identified in June as the leading open-weight model on the benchmark. Between them, K3 and GLM-5.2 mark the current open-weight ceiling on offensive cyber.
- Did Kimi K3 refuse to help with the offensive-cyber tasks?
- No. UK AISI and CAISI report that Kimi K3's built-in safeguards did not prevent it from attempting exploit development or offensive cyber operations during the evaluation. The report frames this as a safety-alignment gap — the model's below-frontier performance is a capability limit, not a refusal.
- Why does a preliminary government evaluation of a Chinese model matter now?
- Kimi K3 shipped 16 July 2026 as Moonshot's 2.8-trillion-parameter frontier open-weight model, and full weights land 27 July. Governments moving first on cyber evaluations before wide download signals that open-weight risk assessment is now a Day-Zero exercise, not a post-hoc one — and gives US policymakers a shared UK-US baseline as sanctions discussions around Moonshot heat up.
Try it
https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities