AI/TLDR

ClawBench Team · 2026-04-09 · notable

ClawBench — 153 Real-World Browser-Agent Tasks Across 144 Live Websites; Best Model at 33.3%

ClawBench evaluates AI browser agents on 153 everyday tasks (purchasing, booking, job apps) across 144 live production websites, using a 5-layer recording pipeline and DOM-match + LLM judge—revealing a 40-point gap vs. sandbox benchmarks.

ClawBench: real-world browser agent benchmark across 144 live websites

Agents score 70% in sandboxes; ClawBench shows the real-world number is 33%

Key specs

GitHub stars154

What is it?

An open benchmark of 153 human tasks—shopping, booking, job applications—spread across 144 real production websites. An interception layer blocks final submissions so evaluation is safe but the website complexity is genuine.

How does it work?

Each task run is captured in 5 layers: screenshots, DOM snapshots, HTTP logs, execution traces, and audit logs. A composite DOM-match + LLM judge produces fine-grained rubric scores across 2,159 checkpoints.

Why does it matter?

Claude Sonnet 4.6 achieves 33.3%. GLM-5 gets 24.2%. No model exceeds 50% in any category. Closes the credibility gap between headline benchmark numbers and what agents can actually do for users.

Try it

https://claw-bench.com

Sources · 3 outlets

Tags

  • dataset
  • benchmark
  • agents
  • browser
  • evaluation
  • web

← All releases · Learn AI