AI/TLDR

UK AISI + CAISI · 2026-07-23 · major

UK AISI + CAISI evaluate Kimi K3 — 32% on ExploitBench vs 76% for US frontier models

UK AISI and NIST's CAISI evaluated Kimi K3 on cyber tasks: 32% on ExploitBench vs 76% for US frontier models, 0 of 41 arbitrary-code-execution outcomes, and step 17 of a 32-step network attack vs 28.5 for the frontier.

UK AISI Kimi K3 cyber capabilities assessment social card

UK AISI and NIST's CAISI put Kimi K3 through joint cyber evals and found it trails the US frontier by a wide margin — but still leads open-weight rivals.

Key specs

Exploit bench (kimi k3)32%
Exploit bench (top us models)76%
Kimi k3 ace (of 41 tasks)0
Network attack steps (kimi k3)17 of 32
Network attack steps (frontier avg)28.5

Quick facts

EvaluatorsUK AISI + US CAISI (NIST)
Model testedKimi K3 (Moonshot AI, released 16 Jul 2026)
Cyber-exploit benchmarkExploitBench (CMU) — 41 Chrome V8 vulnerabilities post-2023
Kimi K3 ExploitBench score32% (vs 76% for top US models)
Arbitrary code execution0 of 41 samples (US frontier avg: 20 of 41)
Simulated network attackstep 17 of 32 avg (US frontier avg: 28.5)
vs GLM-5.2Kimi K3 32% beats GLM-5.2 24% on ExploitBench
Safety alignmentK3 did not refuse offensive-cyber tasks

Benchmarks

ExploitBench (41 Chrome V8 vulns, post-2023)
Kimi K3 (Moonshot)32%
US frontier models (avg)76%
GLM-5.2 (Z.ai)24%
source ↗

What is it?

The UK Artificial Intelligence Security Institute and the US Center for AI Standards and Innovation, part of NIST, published a joint preliminary assessment of Moonshot AI's Kimi K3 on offensive cyber tasks. It is the first cross-Atlantic government evaluation of a Chinese frontier open-weight model on this axis.

How does it work?

The two agencies ran Kimi K3 through Carnegie Mellon's ExploitBench, a set of 41 exploit-development tasks against Chrome V8 vulnerabilities disclosed after 2023, and through a 32-step simulated network-attack range. Each task was scored against the current US frontier closed models and the strongest open-weight model as of June, GLM-5.2, using an Item Response Theory methodology the labs have published in prior reports.

Why does it matter?

The result gives policymakers concrete numbers instead of vibes: Kimi K3 is not at frontier cyber capability, but it is the new open-weight ceiling and its safeguards do not refuse offensive-cyber help. The evaluation lands four days before Moonshot's scheduled 27 July open-weight drop, and while the White House weighs Entity-List designations against Moonshot — giving a shared UK-US technical baseline for those talks.

Who is it for?

Security teams, policy watchers, and open-weight model users trying to size the real cyber-abuse risk of a new frontier Chinese model.

Frequently asked questions

What did UK AISI and CAISI actually measure with Kimi K3?
UK AISI and CAISI ran Kimi K3 through ExploitBench, a Carnegie Mellon University benchmark that asks a model to write exploits for 41 Chrome V8 engine vulnerabilities disclosed after 2023, plus a 32-step simulated network-attack range. They compared its results against the current US frontier closed models and against the strongest open-weight alternatives.
How does Kimi K3 compare to US frontier models on cyber tasks?
Kimi K3 sits well behind. On ExploitBench it scored 32% to the leading US models' 76%. On arbitrary code execution it landed 0 of 41 tasks compared with a US-frontier average of 20 of 41. On the network-attack range it made step 17 of 32 on average, versus 28.5 steps for the strongest models — and completed the full attack in only 1 of 10 attempts.
How does Kimi K3 compare to other open-weight models?
Kimi K3 is now the strongest open-weight model on this cyber suite. Its 32% ExploitBench score beats Z.ai's GLM-5.2 at 24%, which UK AISI had identified in June as the leading open-weight model on the benchmark. Between them, K3 and GLM-5.2 mark the current open-weight ceiling on offensive cyber.
Did Kimi K3 refuse to help with the offensive-cyber tasks?
No. UK AISI and CAISI report that Kimi K3's built-in safeguards did not prevent it from attempting exploit development or offensive cyber operations during the evaluation. The report frames this as a safety-alignment gap — the model's below-frontier performance is a capability limit, not a refusal.
Why does a preliminary government evaluation of a Chinese model matter now?
Kimi K3 shipped 16 July 2026 as Moonshot's 2.8-trillion-parameter frontier open-weight model, and full weights land 27 July. Governments moving first on cyber evaluations before wide download signals that open-weight risk assessment is now a Day-Zero exercise, not a post-hoc one — and gives US policymakers a shared UK-US baseline as sanctions discussions around Moonshot heat up.

Try it

https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities

Sources · 3 outlets

Tags

  • kimi-k3
  • moonshot
  • uk-aisi
  • caisi
  • nist
  • cyber-evaluation
  • exploit-development
  • exploitbench
  • ai-safety
  • government-evaluation
  • open-weight
  • glm-5-2
  • cybersecurity

← All releases · Learn AI