logoalt Hacker News

NanoGPT Speedrun Frontier

38 pointsby staredyesterday at 10:14 PM8 commentsview on HN

Comments

vibe42today at 12:35 AM

"Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind. They preserve weak signals long enough to validate them, but they also have a better understanding of the results."

Curious if a harness that helped preserve signals in some history log would change the outcome.

Also curious if different goal prompts would have changed the outcome. Not a bunch of prompt engineering; small diffs like "consider novel solutions, keep track of weak signals".

IMO they allocated quite a bit of GPU time to the same goal prompt.

ninjahawk1today at 12:20 AM

I might’ve missed it, but why was Fable 5 tested on high while Opus 5 was tested on max? Seems like quite a few of them aren’t on the same effort setting as well. Although effort doesn’t really matter anymore since they can change it dynamically, seems like that might be viewed as an experimental error to some.

skybriantoday at 12:05 AM

Neat!

The graphs show the "best validated result" for each model. I wonder how much variation there is between runs for a model?

totetsutoday at 12:11 AM

“We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.”

Uh.. okay.. but whats a run… read blog

“We want to measure how well frontier models can conduct research….””we ran 153 autonomous runs on the nanoGPT optimizer speedrun across”

Okay but what is a optimiser run and what connection does it have to being good at research?

“For comparison, Anthropic's internal automated AI R&D evaluation optimizes a model on a CPU node,”

So I should go look what Anthropic was doing to understand?

Why not just explain what it means in their blog..

show 1 reply
ninjahawk1today at 12:17 AM

I misread the graph and genuinely thought you put NanoGPT where Fable is.

Lol.

show 1 reply