logoalt Hacker News

westurnertoday at 11:43 AM1 replyview on HN

How does it perform on ARC-AGI-3?

There was this a few weeks ago:

"Schema Harness Achieves ~99% on Arc‑AGI‑3 Public" https://news.ycombinator.com/item?id=48938163

>> Schema, the harness we introduce today, reaches 99% on the ARC_AGI_3 Public set using Claude Opus 4.8 and Fable 5, and 95.35% using GPT‑5.6 Sol

What does that do with 5.6 Luna instead of the expensive models?

What of 'schema' would improve the performance of mdlARC?

mdlARC: https://github.com/mvakde/mdlARC

There's an updated ARC-AGI-1 chart with 5.6 Luna in each thinking level in this video from last week: "A New Architecture [..] | MOONSHOTS " https://youtube.com/watch?v=qQfUbo7Ldc0&t=2m5s


Replies

evilmathkidtoday at 12:07 PM

its not gonna do well on ARC-3 without some significant changes and effort

The new arch in that video is kinda misleading. Didn't really compare against proper baselines