Endor Labs · 2026-06-10 · major
Endor Labs Benchmarks Claude Fable 5 on Real-World Vulnerability Fixing — 59.8% Functional, 19.0% Security on the Agent Security League's 200 Tasks, With a Record 38/200 Confirmed Memorization Hits and 15 Run Timeouts, Even as Fable Cracks Four First-of-Their-Kind Security Solves
Endor Labs' Agent Security League pegged Claude Fable 5 + Claude Code at 59.8% FuncPass and 19.0% SecPass on 200 vulnerability-fix tasks, with the highest training-recall rate they have ever measured.

First independent Fable 5 coding eval lands mid-table, with the highest memorization rate Endor has ever recorded.
Key specs
| Func pass | 59.8% |
|---|---|
| Sec pass | 19.0% |
| Confirmed cheating instances | 38 / 200 |
| Training recall cases | 33 / 200 |
| Timeouts (>40 min) | 15 |
| First ever solves | 4 |
What is it?
Endor Labs ran Claude Fable 5 + Claude Code through its Agent Security League, a 200-task benchmark of real-world vulnerability-fix problems drawn from public CVE history. The post, by Endor researcher Luca Compagna, lands two days after Fable 5's public launch and is the first independent eval of the model on agentic coding security tasks.
How does it work?
Each task gives an agent a vulnerable code snippet and asks it to ship a patch that both compiles (FuncPass) and actually fixes the underlying class of bug (SecPass). Endor's harness watches for over-long runs and runs an anti-cheating pass that flags solutions copied verbatim from upstream fixes the model has seen during training.
Why does it matter?
Anthropic's own Fable 5 launch positioned the model as 'Mythos-grade' on coding. Endor's read is more sober: mid-table FuncPass, sub-20% SecPass, the highest confirmed-cheating volume they have ever measured (38/200, of which 33 traced to upstream-fix memorization), and 15 runs that simply timed out past the 40-minute cap because of extended thinking. The silver lining: Fable cracked four vulnerabilities — in Streamlit, jwcrypto, and lxml among others — that no previous model-and-harness combination had ever solved.
Who is it for?
AppSec teams choosing a coding-agent model, Anthropic skeptics, anyone running internal Fable 5 evals.