AI/TLDR

Artificial Analysis · 2026-09-04 · major

Artificial Analysis Index v4.2 — private test sets now carry 40%

Artificial Analysis Intelligence Index v4.2 retires GPQA Diamond as saturated and adds two evaluations, AA-Briefcase and GDP.pdf. Private held-out test sets now carry 40% of the Index weight, double the figure from v4.1.

Artificial Analysis Intelligence Index v4.2 announcement graphic

The leaderboard most people quote just changed what it measures, and how much of it labs are allowed to see.

Quick facts

Versionv4.2
PublishedSeptember 4, 2026
Evaluations10
AddedAA-Briefcase, GDP.pdf
RetiredGPQA Diamond
Private test weight40%, double v4.1
Top modelClaude Fable 5.1, score 57

What is it?

Two evaluations come in and one goes out in Artificial Analysis Intelligence Index v4.2. AA-Briefcase, an in-house test of realistic agentic knowledge-work tasks in complex projects built by industry experts, joins alongside GDP.pdf from Surge AI. GPQA Diamond leaves the Index after becoming saturated — frontier models now score too closely on it to tell them apart.

How does it work?

GDP.pdf asks a model to reason over 100 PDFs across ten domains, pulling together evidence spread over 4,592 pages of text, tables, charts, footnotes and exclusions in a single turn. AA-Briefcase keeps its test set private. Together the two additions push the private, held-out share of the Index weighting to 40%.

Why does it matter?

A leaderboard labs can train against stops measuring anything. Doubling the private share makes the Index harder to optimize for directly, and Artificial Analysis says the held-out portion will grow again in v5. This is an interim release: the team pulled parts of v5 forward because the frontier moved faster than its normal release cadence.

Who is it for?

anyone comparing frontier models

Frequently asked questions

How do models rank under Artificial Analysis Intelligence Index v4.2?
Under Artificial Analysis Intelligence Index v4.2, Claude Fable 5.1 on adaptive reasoning at max effort leads with a score of 57, GPT-6 Astra at max effort follows at 55, and GPT-6 Astra at xhigh effort takes third at 54. The order flips on the new document evaluation: GPT-6 Astra scores 33.2% on GDP.pdf against Claude Fable 5.1's 26.2%.
Which evaluations make up Artificial Analysis Intelligence Index v4.2?
Artificial Analysis Intelligence Index v4.2 aggregates ten evaluations: AA-Briefcase, GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GDP.pdf, CritPt, AA-Omniscience and AA-LCR v1.1. The suite is primarily text-based and English-language; image input, speech input and multilingual performance are benchmarked separately from the Index.
Is Artificial Analysis Intelligence Index v4.2 the final version?
No — Artificial Analysis calls v4.2 an interim update and says v5 of the Index is already in progress, with more incremental releases planned in the near future. The team pulled elements of v5 forward into v4.2 rather than waiting, because frontier model capability moved faster than its usual release schedule.
Can AI labs see the questions in Artificial Analysis Intelligence Index v4.2?
Not all of them. Private, held-out test sets account for 40% of the weighting in Artificial Analysis Intelligence Index v4.2, double the share in v4.1, and AA-Briefcase in particular keeps its test set private. Artificial Analysis states the goal plainly: make it harder for labs to optimize directly against the benchmark.

Sources

Tags

  • artificial-analysis
  • benchmark
  • leaderboard
  • evaluation
  • intelligence-index
  • aa-briefcase
  • gdp-pdf
  • gpqa-diamond
  • benchmark-saturation
  • held-out-test-set
  • agentic-benchmarks

← All releases · Learn AI