logoalt Hacker News

Ariarule • today at 7:30 PM • 3 replies • view on HN

Always glad to see more open-weight models, but this caption on the 2nd demo image had me do a double-take: "Land or Water Generalization Experiment: We recreated the viral X puzzle by asking Beam to create a fixed 180×90 grid for longitudes -179° to 179° and latitudes -89° to 89°, with 16,200 points. This puzzle is a few days old, so could not appear in the training data, thus testing the model’s generalization. Beam gets 95.5% coverage right, putting us between Opus 5 (92.5%) and Fable 5 (97.8%), which shows how well it generalizes to novel new tasks."

Oof, no, this "puzzle is a few days old" is incorrect even if it's a social media trend just recently. Asking a model to generate a world map in this way is _at least_ from August 2025 as it appeared on LessWrong at that time: https://www.lesswrong.com/posts/xwdRzJxyqFqgXTWbH/how-does-a...


Replies

extr • today at 7:48 PM

Yeah I remember when the original post about this came out. Def not recent. Though I think their point survives in that they didn't exactly RL on this.

jiggawatts • today at 10:13 PM

I'd love to see these tests repeated for the current frontier models...

charlieyu1 • today at 8:49 PM

I don't think age of the puzzle even matters, all models have search capacities these days

➕ show 3 replies