logoalt Hacker News

pvillanotoday at 5:33 PM5 repliesview on HN

Imagine yourself the CEO of a big AI company. It takes about a month to develop and train a model, so you release a new model every month. A startup says they can 10x your efficiency. What does that get you? You can't release a new model every three days. You can't 10x R&D either. You definitely can't tell investors that you are growing at the same rate, but selling off assets and cancelling purchasing contracts. So you just never improve efficiency enough to use less energy than the previous model version.

I don't believe this is actually happening.


Replies

alex_duftoday at 6:40 PM

A 10x reduction in pre-training means a 10x faster feedback loop. I'm pretty sure any lab would sign up for that. You can start experimenting on different approaches much more aggressively.

impossibleforktoday at 8:22 PM

So you can make a huge internal model that you can then distill from?

enzyme1234today at 6:02 PM

the obvious answer is that you can try more experiments over the same period of time, so you find more improvements per month, and the rate of improvement of the models you release increases

awestroketoday at 5:40 PM

> What does that get you?

Cheaper model training runs? Ability to scale training to larger model sizes without extending training time?

show 1 reply
BoorishBearstoday at 6:00 PM

This seems weirdly pessemistic: frontier labs have much stronger pretraining than most open weights models

And reading the release it feels very obvious this is also a ton of aligning their data mix with coding and science: we don't know that this model doesn't have terrible world knowledge or is ruined for anything related to subjective preference

They also repeatedly mention knowledge almost as if they saw that skepticism coming, but then limit knowledge to topics where more understanding of how code/scientific writing looks would produce the same graph as having actual world knowledge maintained.

That's not nefarious (they literally build coding models), but it also means the resulting model isn't necessarily competitive with a frontier model in a broader way.

This feels like the inverse approach to what Thinking Machines did with Inkling (trying to train as "un-spikey" a base model as possible)