HeadlinesBriefing favicon HeadlinesBriefing.com

Brood War Bench: AI Models Tested in StarCraft

Hacker News •
×

Brood War Bench tested AI models at StarCraft: Brood War. None played beyond beginner level. Codex Astra was the clear leader, beating all other models consistently. Grok models were not smart enough to play yet.

Older models treated RTS as turn-based, getting destroyed while thinking. Newer models were more cognizant of thinking costs, though some lower-effort settings performed better.

Codex's strongest recurring idea was disruption, sending Probes to attack workers. It often created separate subagents for economy, army production, and control that didn't communicate well. Grok 4.6 spent long stretches reasoning with few command batches — one run logged 11,138 tokens but issued only six batches in 43 minutes with no combat unit.

Claude Fable earnestly tried to play, building economies and climbing tech trees. It reached Mutalisks, Lairs, Spires, and advanced tech buildings before winning some games.