AI/TLDR

Ben Swerdlow · 2026-09-19 · major

Brood War Bench — Codex Astra wins all 18 of its StarCraft games

Brood War Bench puts 19 AI agent setups into StarCraft: Brood War and has each one play every other. Codex Astra at xhigh effort won all 18 of its games. Grok 4.6 and Claude Haiku won almost none.

Brood War Bench report card for an AI agent StarCraft tournament

A round-robin StarCraft: Brood War tournament where 19 AI agent setups play each other instead of solving coding tasks.

Key specs

Agent setups19
Best record18-0

Quick facts

What it testsReal-time StarCraft: Brood War play by LLM agents
FormatRound robin — every setup plays every other
Models comparedCodex Astra, Codex 5.6, Claude, Grok 4.6
WinnerCodex Astra / xhigh — 18-0, 100%
Cost per game$0.16 – $21.07
Runs onFreestyle VMs, matches in parallel

Benchmarks

Brood War Bench win rate
Codex Astra / xhigh100%
Codex Astra / medium88.9%
Claude Fable83.3%
Codex 5.6 Sol / medium72.2%
Claude Opus 566.7%
Claude Sonnet38.9%
Grok 4.6 / xhigh11.1%
Claude Haiku0%
source ↗

What is it?

Brood War Bench scores AI agents on a real-time strategy game rather than a coding or reasoning test. Ben Swerdlow built a version of StarCraft: Brood War that can only be played through agents, then ran a round-robin in which every model and effort setting faced every other one. The leaderboard lists 19 setups with wins, losses, actions per minute and dollar cost per game.

How does it work?

Matchups ran in parallel on Freestyle virtual machines, which saved the game-engine data and both agents' harness logs for every game. The clock never stops while a model thinks, so slow reasoning is punished on the spot: in one game Grok 4.6 logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit. Codex went the other way and spawned separate subagents for economy, army production and army control, though the report says they rarely coordinated well.

Why does it matter?

Most agent evaluations let a model take as long as it likes on each step. Real-time play removes that cushion, so Brood War Bench measures how well a model converts thinking into timely action — and the cost column prices that thinking, from $0.16 to $21.07 a game. Swerdlow adds that no agent played above beginner level, which leaves plenty of room to improve.

Who is it for?

agent researchers and eval builders

Frequently asked questions

Which setup won Brood War Bench?
Codex Astra at xhigh reasoning effort won Brood War Bench with an 18-0 record and a 100% win rate. Codex Astra at medium effort came second at 88.9%, and Claude Fable placed third with 15 wins and 3 losses for 83.3%. The report notes that Codex often won by sending a worker across the map to harass the opponent early.
How much does one Brood War Bench game cost?
Cost per game in Brood War Bench ranges from $0.16 for Codex 5.6 Luna at xhigh effort up to $21.07 for Codex Astra at low effort. The winning Codex Astra xhigh setup averaged $10.54 a game, while Claude Opus 5 averaged $20.78 and Claude Haiku averaged $0.34. Cheap does not mean weak, and expensive does not mean strong.
How did Grok 4.6 do compared with Codex Astra?
Grok 4.6 finished near the bottom of Brood War Bench in all three effort settings: 11.1% at xhigh, 5.6% at medium and 0% at low. Codex Astra took the top three spots by contrast. Swerdlow's report attributes the gap to Grok producing long stretches of reasoning and very few command batches, which loses games in a real-time setting.
Does more reasoning effort help an agent play StarCraft better?
Not consistently, on Brood War Bench numbers. Codex Astra improves with effort — 77.8% at low, 88.9% at medium, 100% at xhigh — but Codex 5.6 Sol moves the other way, scoring 72.2% at medium and only 61.1% at xhigh. Codex 5.6 Luna is worst at medium effort. Thinking longer costs game time, so it does not always pay off.

Sources · 2 outlets

Tags

  • benchmark
  • agents
  • evaluation
  • starcraft
  • games
  • real-time-strategy
  • codex
  • claude
  • grok
  • llm-agents

← All releases · Learn AI