Ben Swerdlow · 2026-09-19 · major
Brood War Bench — Codex Astra wins all 18 of its StarCraft games
Brood War Bench puts 19 AI agent setups into StarCraft: Brood War and has each one play every other. Codex Astra at xhigh effort won all 18 of its games. Grok 4.6 and Claude Haiku won almost none.

A round-robin StarCraft: Brood War tournament where 19 AI agent setups play each other instead of solving coding tasks.
Key specs
| Agent setups | 19 |
|---|---|
| Best record | 18-0 |
Quick facts
| What it tests | Real-time StarCraft: Brood War play by LLM agents |
|---|---|
| Format | Round robin — every setup plays every other |
| Models compared | Codex Astra, Codex 5.6, Claude, Grok 4.6 |
| Winner | Codex Astra / xhigh — 18-0, 100% |
| Cost per game | $0.16 – $21.07 |
| Runs on | Freestyle VMs, matches in parallel |
Benchmarks
| Codex Astra / xhigh | 100% | |
|---|---|---|
| Codex Astra / medium | 88.9% | |
| Claude Fable | 83.3% | |
| Codex 5.6 Sol / medium | 72.2% | |
| Claude Opus 5 | 66.7% | |
| Claude Sonnet | 38.9% | |
| Grok 4.6 / xhigh | 11.1% | |
| Claude Haiku | 0% |
What is it?
Brood War Bench scores AI agents on a real-time strategy game rather than a coding or reasoning test. Ben Swerdlow built a version of StarCraft: Brood War that can only be played through agents, then ran a round-robin in which every model and effort setting faced every other one. The leaderboard lists 19 setups with wins, losses, actions per minute and dollar cost per game.
How does it work?
Matchups ran in parallel on Freestyle virtual machines, which saved the game-engine data and both agents' harness logs for every game. The clock never stops while a model thinks, so slow reasoning is punished on the spot: in one game Grok 4.6 logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit. Codex went the other way and spawned separate subagents for economy, army production and army control, though the report says they rarely coordinated well.
Why does it matter?
Most agent evaluations let a model take as long as it likes on each step. Real-time play removes that cushion, so Brood War Bench measures how well a model converts thinking into timely action — and the cost column prices that thinking, from $0.16 to $21.07 a game. Swerdlow adds that no agent played above beginner level, which leaves plenty of room to improve.
Who is it for?
agent researchers and eval builders
Frequently asked questions
- Which setup won Brood War Bench?
- Codex Astra at xhigh reasoning effort won Brood War Bench with an 18-0 record and a 100% win rate. Codex Astra at medium effort came second at 88.9%, and Claude Fable placed third with 15 wins and 3 losses for 83.3%. The report notes that Codex often won by sending a worker across the map to harass the opponent early.
- How much does one Brood War Bench game cost?
- Cost per game in Brood War Bench ranges from $0.16 for Codex 5.6 Luna at xhigh effort up to $21.07 for Codex Astra at low effort. The winning Codex Astra xhigh setup averaged $10.54 a game, while Claude Opus 5 averaged $20.78 and Claude Haiku averaged $0.34. Cheap does not mean weak, and expensive does not mean strong.
- How did Grok 4.6 do compared with Codex Astra?
- Grok 4.6 finished near the bottom of Brood War Bench in all three effort settings: 11.1% at xhigh, 5.6% at medium and 0% at low. Codex Astra took the top three spots by contrast. Swerdlow's report attributes the gap to Grok producing long stretches of reasoning and very few command batches, which loses games in a real-time setting.
- Does more reasoning effort help an agent play StarCraft better?
- Not consistently, on Brood War Bench numbers. Codex Astra improves with effort — 77.8% at low, 88.9% at medium, 100% at xhigh — but Codex 5.6 Sol moves the other way, scoring 72.2% at medium and only 61.1% at xhigh. Codex 5.6 Luna is worst at medium effort. Thinking longer costs game time, so it does not always pay off.