logoalt Hacker News

CamperBob2yesterday at 7:29 PM1 replyview on HN

At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.


Replies

jaggederestyesterday at 7:54 PM

I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)