logoalt Hacker News

vlovich123 • yesterday at 9:54 PM • 1 reply • view on HN

> The "does not rescue it". No human would write like that.

This is what I don’t understand. Supposedly LLMs are trained on human text. Why do they come up with such unrealistic prose? Is it intentional because the companies want the tells to be obvious?


Replies

mediaman • yesterday at 11:01 PM

They’re not just trained on human prose. They’re sent to RLHF, and also their language changes as a result of RL on verifiable rewards.

Getting it to write well is really hard because there’s no real way to verify whether it’s good prose or not. You and I can tell, but we can’t write a verifier that codifies our judgment.

Maybe they’ll find a way to improve this, but for now it’s certainly one of the harder problems to solve for LLMs.

Part of it is that I think they also have poor theory of mind, which I imagine is also a hard thing to train it to do.

➕ show 1 reply