NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute

94 points • by sdpmas • yesterday at 6:51 PM • 17 comments • view on HN

Comments

What's the human baseline? How many cats does a human need to see to learn what a cat is, vs an AI?

Maybe not quite a fair comparison since my human brain has been "learning" for half a billion years before I was born.

I wonder if there's an equivalent of that for AI. Evolving the architectures?

➕ show 2 replies

In their little algorithm box on Chain Distillation, they have at step 2b some expression that involves multiplying and dividing by `T`, and then they say "where α = 0.5, T = 1.0".

I think someone during the copy-editing process told them this needed to look more complicated?

➕ show 1 reply

nsnzjznzbx • yesterday at 8:59 PM

We will get to the point where you can quickly bootstrap i.e. an LLM can train a better LLM in a loop, leave it and it can really learn. Like learn learn.

"Train yourself to solve this problem see OBJECTIVE.md"

➕ show 1 reply

1425curlz80 • today at 12:31 AM

the comments here are better than the article lol

littlestymaar • yesterday at 7:38 PM

> Data efficiency matters because compute grows much faster than data [2] (referencing a paper from 2022)

I'm not convinced this is particularly true in today's world, if you have more compute, you can simply generate more, and higher quality, artificial data. That's what all labs have been doing since at least 2023.

Also, the post references the Chinchilla-optimal training as a comparison baseline, but everyone has moved far beyond Chinchilla scaling, small models are routinely trained on 10-400 times more data than (1-40T tokens) than the Chinchilla-optimal number, so the entire industry went the complete opposite of what they are proposing.

That doesn't mean the techniques presented here are useless or anything (I'm not qualified to judge) but you should take the introduction with a grain of salt.

➕ show 3 replies

yorwba • yesterday at 7:30 PM

Related: Discussion on the initial NanoGPT Slowrun announcement: https://news.ycombinator.com/item?id=47251259 (185 points 15 days ago, 39 comments)

➕ show 1 reply

myylogic • yesterday at 9:46 PM

[dead]

aledevv • yesterday at 10:32 PM

[dead]

AliEveryHour16 • yesterday at 9:04 PM

[dead]

alt Hacker News

NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute

Comments