logoalt Hacker News

lelanthranyesterday at 2:22 PM1 replyview on HN

> It takes more computing power, and money, to pay for training an LLM, compared to distilling an LLM that someone else has already trained.

And that is still less money than it took to create that data in the first place, which the AI companies then gladly took to use for training.


Replies

rmunnyesterday at 11:51 PM

Which is true, but also completely irrelevant to the point of this discussion. Which was about the U.S. companies spending trillions and the Chinese companies spending billions and matching them in quality. I made the point that the U.S. companies and Chinese companies were doing different things: training on raw data vs. distilling the model that someone else had trained. That difference explains the difference in spending. (The secondary point is that the Chinese companies would not have been able to get to the point they have, while spending as little as they have, without Claude, GPT, et al to copy from).

This is probably the last reply I'll make to you. I'm getting a little tired of repeating myself. Whether you're just not getting it, or refusing to get it, either way it's starting to feel like a waste of time to try to rephrase the same thing again and again. Please read more carefully in the future.