Brood War Bench

293 points · 127 comments on HN · read original →

Points and comments are a snapshot, not live.

AI agents including Codex, Claude, and Grok play StarCraft: Brood War at beginner level; Codex Astra is strongest.

Ben Swerdlow built a StarCraft: Brood War benchmark where AI agents play head-to-head via a harness. Codex Astra (xhigh) won 18-0 (100%), followed by Codex Astra medium at 88.9% and Claude Fable at 83.3%. Grok 4.6 could not play effectively, often issuing only a few command batches per game. Even the best agents showed beginner-level play: poor macro, weak defense, and no complex strategies. Codex relied on early harassment (probe rushes) but struggled with sustained production. Fable attempted tech progression but failed to execute. Cost per game ranged from $0.42 (Codex Luna low) to $21.07 (Codex Astra low).

What commenters are saying

Commenters praised the benchmark as a novel test of long-term strategy and tactical reasoning for LLMs. Several shared personal nostalgia for StarCraft: Brood War and its community. One noted that Zerg is strongest at high levels due to production scaling and defilers. Another recommended GoBench (9x9 Go) as a similar strategy benchmark. A commenter suggested LLMs could improve by analyzing their own games and writing instructions for future matches. The harness used Claude Code, Codex, and Grok Build for cost reasons, with BW-API for commands and observations. Some questioned whether training data includes game transcripts, affecting results.