I was thinking about this the other day. If we did a plot of 'model ability' vs 'comp...

woggy • yesterday at 8:56 PM • 2 replies • view on HN

I was thinking about this the other day. If we did a plot of 'model ability' vs 'computational resources' what kind of relationship would we see? Is the improvement due to algorithmic improvements or just more and more hardware?

Replies

chasd00 • yesterday at 9:51 PM

i don't think adding more hardware does anything except increase performance scaling. I think most improvement gains are made through specialized training (RL) after the base training is done. I suppose more GPU RAM means a larger model is feasible, so in that case more hardware could mean a better model. I get the feeling all the datacenters being proposed are there to either serve the API or create and train various specialized models from a base general one.

ryoshu • yesterday at 9:07 PM

I think the harnesses are responsible for a lot of recent gains.

➕ show 1 reply

alt Hacker News

Replies