logoalt Hacker News

kzrdudetoday at 7:14 AM1 replyview on HN

We need to recognize this as a failure in training. It did some useful stuff but it can be much better. A training signal is likely missing.


Replies

NiloCKtoday at 7:25 AM

As I understand it, the rough guess as to what's happening here is that most recent capabilities progress comes from specific verifiable-rewards reinforcement training (RL). The RL pressures are all about task performance, but (surprise surprise) highly human-legible English language usage isn't very important to the models abilities to address the tasks.

Weirdly enough, the pressures are having them drift toward novel dialects of English that work well for their own chains of thought. Open question about whether they'd drift all the way to a new language given enough time.

show 1 reply