HeadlinesBriefing favicon HeadlinesBriefing.com

Pac-Bench: One-Shot Pac-Man Model Benchmark

Hacker News •
×

Hacker News user jonclegg launched Pac-Bench to evaluate how effectively AI models can generate a playable Pac-Man clone from a single prompt. The challenge requires models to create a complete HTML page in one attempt, with no follow-up prompts or corrections allowed. Results reveal significant variation in performance, with some models producing functional games while others struggle with basic mechanics.

The benchmark tracks success rates, code efficiency, and adherence to the strict one-shot constraint, providing insights into current AI coding capabilities. Organizers collected data on token usage, generation time, and HTML output size to measure practical viability. The project page hosts individual model attempts and detailed metrics for community analysis.