Pangram verdict · v3.3
We believe that this text is a mix of AI and human-written content.
AI likelihood · overall
AIArticle text · 300 words · 2 segments analyzed
Six of them played 100 games each against the arcade ghosts. See how they rank, watch their games or play against them. jev 1.130 High score0 GhostsClassic Watch Leaderboard 100 games per model against the classic ghosts. =1 means tied: those scores are within each other's margin of error.
Benchmark your own model jevman is open source, and any model behind an HTTP endpoint can play it: a hosted model, a fine-tune or one running on your laptop. 01 Put your model behind an endpoint At every junction the game sends the maze as JSON and asks for a direction, with odds if your model has them. The repo has a 34-line example endpoint to start from. 02 Run the benchmark One command plays the games under the leaderboard's rules: real time, 2 seconds per answer, the classic ghosts, three lives or 5 minutes. 03 Send a pull request CI replays every game you submit and checks its score. Your model then joins the leaderboard as self-reported. Method GamesEach model played 100 games against the classic scripted ghosts, each until it lost all three lives. Games are capped at 5 minutes so every run ends, but none came close: the longest lasted 2 minutes 24 seconds. DecisionsAt every junction the game asks the model one question, with the maze, pellets and ghosts as state. The model returns a probability per direction, and Pac-Man takes its pick. DeadlineAn answer that takes longer than 2 seconds is replaced by a simple backup rule, counted under backup moves. ScoresThe ranking uses the mean score with a 95% margin of error (±2 standard errors), and models within each other's margin are tied. The high score is a model's best single game. ChecksEvery game is recorded and replays exactly, so any result can be checked.