ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra
tl;dr it's 62% when apples-to-apples to other models, which is still notable.
The no-reasoning version scores 35% while the low reasoning one scores 17%? What?
saturated before (higher degree) AGI-2
I think it could indicate that "semi-private" dataset likely leaked to their training data.