logoalt Hacker News

Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?

67 points • by thefourthchime • yesterday at 10:43 PM • 42 comments • view on HN

Benchmarks how well Harness+models can create a Pac-Man game from a single prompt:

“Create a Pac-Man game in a single HTML page”

Each model gets one shot — no follow-up prompts or fixes.


Comments

yambam • today at 7:03 AM

This reminds me of a couple decades ago when some friends and I set ourselves a challenge to, individually, each create as much of a Pac-Man clone as possible in 10 hours. None of us had any experience in games programming or graphics coding. It was great fun and we all learned a lot.

No-one ended up with a complete clone but I loved how we all ended up focusing on different things, like pixel-perfect graphics versus accuracy in gameplay, and how we all brought our existing skills to the challenge despite not really knowing what we were doing.

I expect if we had AI models available it would have ruined the pleasure of figuring it out for ourselves. I feel kind of sad for the next generation of developers who won't have that experience.

➕ show 1 reply
drcxd • today at 8:04 AM

Interesting, recently I am working on my own clone of Pac-Man. LLM implementations lose lots of details. They are not 1:1 replication of the original game. For example, the behavior of the ghost is not the same as the original. If I have not implemented the game myself, I can not tell the differences. What LLM produced look like the original game, but they are not.

strataspace • today at 12:38 AM

I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering.

The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.

➕ show 3 replies
jmathai • today at 5:14 AM

This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving.

Remember when people considered you a genius for prompting with "You are a skilled writer....".

➕ show 1 reply
xnx • today at 10:42 AM

This would be a much better benchmark if it contained some twist which was not in the training set (e.g. pelicans on bicycles were not common before simonw).

Make a pacman game where pacman can always eat ghosts, but ghosts drop pellets.

_matthew_ • today at 4:57 AM

I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.

continuational • today at 5:11 AM

Here's GLM-5.2: http://show.ahnfelt.net/demos/glm5.2-pacman.html

➕ show 1 reply
swingboy • today at 9:30 AM

Interesting the `high` effort on the Claude models. Did you find that going above that is unnecessary?

nedo_var • today at 4:46 AM

Curious if models grasp the ghost patterns or just react. Pac-Man's more complex than it seems for one-shot learning.

weitendorf • today at 7:04 AM

Gonna be rude and say I don't think this is an interesting or useful benchmark tbh.

Clearly some new RLAAS/dataset/env is being used for this now (it doesn't even seem that complicated, you have one LLM judge whether gameplay is recognizable as the original game or not and another trying to implement a logically/semantically identical version of the game). It's why the performance improvement on this workload has been so dramatic.

Everything is going to go from 0->1 on this benchmark in short order because of that.

mrblinky • today at 6:57 AM

Pretty underwhelming. Surely they’ve ingested some existing code on the web and regurgitating that.

➕ show 1 reply
mikojan • today at 7:36 AM

Clearly the high-tech plagiarism machine will successfully plagiarise one of the most plagiarised games.

hoistway • today at 4:51 AM

Always assumed Pac-Man was an easy solve for modern AI. Guess those ghost patterns are trickier than they look for one-shot learning.

➕ show 2 replies
pranavmore • today at 7:33 AM

Tried opus 5.5 for the same?

Computer0 • yesterday at 10:46 PM

Opus 5-5 seemed like a perfect clone, with others displaying flaws in an initial look. Astra notably created a bunch of surrounding ugly crap to look at.

➕ show 2 replies
Tanjreeve • today at 8:51 AM

I thought this was about playing a pacman game. I feel like we already established that LLMs can create simple games and "local" apps at this point

blindflag • today at 4:57 AM

I'm curious; can you explain why you picked Pac-Man, in particular?

➕ show 1 reply
Stitch4223 • today at 6:07 AM

Someday we’re gonna miss the times where models create jank.

These results get better and better, but the half-baked pac mans games and pelicans are just funny.

You cannot prompt that level of brokenness. Same for those slop-posters you see everywhere.

tamimio • today at 6:40 AM

It would be great to do the models tests but with different harnesses too to see how they impact on the same models and efforts.

diamondDrill • today at 8:40 AM

coolio thanks

alescalaios • today at 10:00 AM

[dead]