Hacker News — AI on Front Page · · 3 min read

Brood War Bench

Mirrored from Hacker News — AI on Front Page for archival readability. Support the source by reading on the original site.

201 pts · 81 comments on Hacker News

Key takeaways

  • None of the models played beyond a beginner level.
  • Codex Astra is the clear leader beating all other models consistently.
  • Grok models are not smart enough to play Brood War yet.
  • Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. Newer models sometimes fell into the same trap, which may explain why some lower-effort settings performed better, but overall were much more cognizant of the cost of thinking.

Powered By Freestyle

Leaderboard

RankSystemWinsLossesAPMCost / gameWin rate
🥇
Codex Astra / xhigh
18012.6$10.54100.0%
Codex Astra / medium
16217.2$15.1188.9%
Claude Fable
15312.6$12.2483.3%
Codex Astra / low
14425.7$21.0777.8%
Codex 5.6 Sol / medium
13510.1$5.1272.2%
Codex 5.6 Sol / low
12618.1$9.2366.7%
Claude Opus 5
12610.5$20.7866.7%
Codex 5.6 Sol / xhigh
1178.0$3.2361.1%
Codex 5.6 Luna / low
9923.8$0.4250.0%
Codex 5.6 Terra / xhigh
9915.8$2.1050.0%
Codex 5.6 Terra / medium
81010.5$3.1544.4%
Codex 5.6 Terra / low
81048.3$4.6544.4%
Codex 5.6 Luna / xhigh
7115.2$0.1638.9%
Claude Sonnet
7116.2$8.9838.9%
Codex 5.6 Luna / medium
61214.8$0.3033.3%
Grok 4.6 / xhigh
2152.8$0.6611.1%
Grok 4.6 / medium
1163.2$0.795.6%
Claude Haiku
0160.3$0.340.0%
Grok 4.6 / low
0164.2$1.260.0%

Brood War Bench started after I built a version of Brood War that you could only play through agents as an experiment to play with friends. I played it with a couple friends who did surprisingly well for people who have only played a couple Starcraft games in their lives. When I asked them why, they said they hadn't done much, they asked their agent to attack and it had built a small army and done the full attack for them. This lead me to wonder how far they can go on their own; this is my answer.

What I observed

01

Codex found cheese before it found macro

Codex's strongest recurring idea was disruption. In Protoss games it often sent a Probe across the map to attack workers or buildings. This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else.

The same systems were much weaker at sustained production. They delayed tech, trickled one or two basic units into defended bases, and threw workers into last stands.

I also noticed Codex often created separate subagents to manage the economy, army production, and army control. They didn't communicate much with one another, so the army agent often sent each new unit straight into an attack, unaware of the larger army the other agents were planning to build.

This is a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing. In games where I helped direct Codex, it was much better at planning those moments and getting its subagents to work together.

The persistence was real. In G009, after losing its army and main base, Codex 5.6 Terra / medium lifted its last Command Center and moved it toward the opposite corner. It survived for another six minutes.

Six Probes cross the map
A Probe first, then Zealots in drips
The last Command Center runs

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from Hacker News — AI on Front Page