logoalt Hacker News

edottoday at 2:09 AM0 repliesview on HN

Right, which was true at the time. So hundreds of billions of dollars have been poured into making LLMs better at these tasks via pretraining, RL, RLHF, post training, etc. again all with something verifiable in the loop. In order to improve the thing in the loop, the loop itself needs to be verifiable.

There have only been a few thousand wars, and they’re all different and all different in the world in which they occurred. The dimensionality is absurd, which is not a problem for LLMs if there’s enough data, but in this case there isn’t.