Anthropic · 2026-08-12 · notable
Conceptual Reasoning Index — scoring models on questions with no right answer
The Conceptual Reasoning Index (CRI) scores language models on questions that cannot be marked right or wrong, such as AI risk, philosophy and decision theory. Opus 5 leads at 73.6 against an estimated ceiling near 91.

One number for how well a model reasons about questions that have no checkable right answer.
Key specs
| Top score (opus 5) | 73.6 |
|---|---|
| Estimated ceiling | 91 |
What is it?
The Conceptual Reasoning Index (CRI) targets a gap in AI evaluation: questions about advanced AI risk, philosophy and decision theory where no grader can mark an answer correct. Redwood Research built the index with Anthropic and published it on Anthropic's alignment science blog on 12 August 2026. Scores run from 0 to 100, and the authors put the practical ceiling at roughly 91.
How does it work?
Three separate benchmarks are combined into one weighted average. LMCA carries 60% and asks a model to judge the quality of conceptual arguments, drawing on 560 position texts and 1,461 expert-rated arguments written against them. ACCoRD carries 20% and checks whether a model stays consistent across related questions, using about 14,000 constraints spread over 18 constraint types. DTBench, 407 decision-theory multiple-choice questions, supplies the final 20%.
Why does it matter?
Most benchmarks reward work that comes with a grading key — maths, code, factual recall. The Conceptual Reasoning Index instead measures the judgment needed to help people think through advanced AI risk, where no key exists. Opus 5 tops the current list at 73.6, well below the estimated 91 ceiling, which gives safety researchers a concrete gap to track as models improve.
Who is it for?
alignment and model-evaluation researchers
Try it
https://conceptualreasoning.ai