logoalt Hacker News

saberienceyesterday at 6:31 PM1 replyview on HN

Arc-AGI (and Arc-AGI-2) is the most overhyped benchmark around though.

It's completely misnamed. It should be called useless visual puzzle benchmark 2.

It's a visual puzzle, making it way easier for humans than for models trained on text firstly. Secondly, it's not really that obvious or easy for humans to solve themselves!

So the idea that if an AI can solve "Arc-AGI" or "Arc-AGI-2" it's super smart or even "AGI" is frankly ridiculous. It's a puzzle that means nothing basically, other than the models can now solve "Arc-AGI"


Replies

CuriouslyCyesterday at 6:38 PM

The puzzles are calibrated for human solve rates, but otherwise I agree.

show 1 reply