OpenAI has just launched its decisions endpoint, and Cloudflare recently released Clef, with many more jev alternatives now available. To compare the popular options, the Opper AI team decided to test them with Pac-Man, a simple benchmark for fast decision making.
The team let jev 1.13, kev, clef, clef flash, GPT-6 Luna and Laya play Pac-Man against classic bot ghosts. The low latency of these models allows for real-time play. Each model played 100 games, and the results were published on a leaderboard. The repository is open source, so anyone can run their own model and join the ranking.
Each game costs about 2 cents, and all models run through Opper, the team's startup. Visitors can also play as Pac-Man themselves, with the ghosts controlled by a mix of models or by a single model such as jev. Free credits are provided so everyone can try it.
Any model behind an HTTP endpoint can take part, whether it is hosted, fine-tuned or running on a laptop. Each game is recorded and replays exactly, so results can be verified. Scores use the mean with a 95% margin of error, and the current leader, jev 1.13, holds the high score. The team is seeking feedback from the community.
Source: Hacker News · Summarized by HeadlinesBriefing