Pac-Bench
Open SourceBenchmark evaluating how well AI models one-shot a Pac-Man game
Pac-Bench is an interactive evaluation framework designed to test the code generation and spatial reasoning capabilities of LLMs by challenging them to one-shot build a fully working Pac-Man game. It provides developers and researchers with a standardized benchmark to measure frontier model performance on complete game development tasks.
Challenges AI models to write a complete, playable Pac-Man game in a single prompt.
Compares performance across various frontier LLMs to gauge coding and spatial reasoning.
Allows users to view and test the resulting generated games directly in the browser.
AI researchers, prompt engineers, and developers interested in evaluating LLM code generation capabilities.
Compare other freshly discovered tools and open-source alternatives.