There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete ...

andrepd • yesterday at 3:32 PM • 3 replies • view on HN

There was a well-publicised "Claude plays Pokémon" stream where Claude failed to complete Pokemon Blue in spectacular fashion, despite weeks of trying. I think only a very gullible person would assume that future LLMs didn't specifically bake this into their training, as they do for popular benchmarks or for penguins riding a bike.

Replies

dwaltrip • yesterday at 7:40 PM

If they game the pelican benchmark, it’d be pretty obvious.

Just try other random, non-realistic things like “a giraffe walking a tightrope”, “a car sitting at a cafe eating a pizza”, etc.

If the results are dramatically different, then they gamed it. If they are similar in quality, then they probably didn’t.

➕ show 1 reply

ctoth • yesterday at 6:04 PM

> as they do for popular benchmarks or for penguins riding a bike.

Citation?

➕ show 1 reply

criley2 • yesterday at 3:46 PM

While it is true that model makers are increasingly trying to game benchmarks, it's also true that benchmark-chasing is lowering model quality. GPT 5, 5.1 and 5.2 have been nearly universally panned by almost every class of user, despite being a benchmark monster. In fact, the more OpenAI tries to benchmark-max, the worse their models seem to get.

➕ show 3 replies

alt Hacker News

Replies