logoalt Hacker News

totetsutoday at 12:11 AM1 replyview on HN

“We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.”

Uh.. okay.. but whats a run… read blog

“We want to measure how well frontier models can conduct research….””we ran 153 autonomous runs on the nanoGPT optimizer speedrun across”

Okay but what is a optimiser run and what connection does it have to being good at research?

“For comparison, Anthropic's internal automated AI R&D evaluation optimizes a model on a CPU node,”

So I should go look what Anthropic was doing to understand?

Why not just explain what it means in their blog..


Replies

deractoday at 12:20 AM

I think that's explained here:

https://www.primeintellect.ai/blog/measuring-autonomous-rese...

Basically they do 8 runs trying to optimize to under 3.28 loss in the fewest training steps possible under time/token constraint. I dunno why 18 * 8 != 153 (it's 144)