Princeton University · 2026-04-23 · major
HAL — Princeton's Holistic Agent Leaderboard Accepted at ICLR 2026, Now Tracks 26K+ Rollouts
The Holistic Agent Leaderboard, accepted at ICLR 2026, provides standardized cost-aware agent evaluation across 9 benchmarks covering coding, web navigation, science tasks, and customer service. It now tracks 26,597 rollouts and has shifted focus to multi-dimensional reliability measurement.

HAL standardizes agent evaluation across 9 benchmarks with cost-tracking — revealing 100x cost differentials for 1% accuracy gains.
What is it?
The Holistic Agent Leaderboard (HAL) from Princeton's SAgE team was accepted as a conference paper at ICLR 2026 and presented at the Rio de Janeiro conference in late April. It provides a unified, open-source harness for reproducible, cost-controlled agent benchmarking across multiple real-world task domains.
How does it work?
HAL orchestrates parallel evaluations across hundreds of VMs, reducing evaluation time from weeks to hours. It covers 9 distinct benchmarks spanning web assistance (AssistantBench, GAIA, Online Mind2Web), scientific programming (CORE-Bench, Scicode, ScienceAgentBench), software engineering (SWE-bench Verified Mini), customer service (TAU-bench Airline), and programming competitions (USACO). Pareto frontier visualizations reveal cost-performance tradeoffs across agents.
Why does it matter?
HAL's key finding is that agents can be 100x more expensive while only 1% better on accuracy, which leaderboards without cost-reporting completely hide. The platform has also shifted focus to reliability measurement — showing that agent performance can drop from 60% to 25% under multi-run consistency checks. With 26,597 rollouts logged, it is becoming the reference infrastructure for third-party agent evaluations.