It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/
I have played few of them and it seems that Opus 5.5 is the first one who really made playable PacMan clone game. On mobile as well.
Could it be because the model was somehow pre-trained? If we compare it with pelicans that are still not-perfect…
Cool page and benchmark idea! Would be nice if there was some kind of grading the results, maybe on different criteria (aesthetic, implementation complexity, correctness, ...). Of course as a one-shot and greenfield benchmark the results are not indicative for all kinds of usage patterns. But as some sibling said, maybe they can be indicative on some general characteristics (especially since the task is so open-ended).
Very cool. I'd love to see someone with access to plenty of token$ make something similar for the "Browser Desktop OS" test. That seems like a pretty comprehensive test thats also fun to test just like this!
Thank you for doing this. It is very helpful not just for capabilities but also for costs.
GPT models really have no taste huh.
I was able to get similar with Qwen 3.8 27B with one shot. I think this game is too well in the training data.
Interesting. Sonnet 5 was horrible, and Opus 5 was unplayable, but both Sonnet 5.5 and Opus 5.5 were about as close to the real thing.
oh! How about pengo, dig dug, & defender?
Wow, that "bake off" page is better than any coding benchmark I've seen! You can really sense the strengths and weaknesses of each model/harness combo.
Oh, it coded a Pac-Man clone. The clone was so good that I thought it was premade in some way and that Sonnet was going to play PacMan.