logoalt Hacker News

red75prime • today at 6:24 AM • 0 replies • view on HN

This is simplistic to the point of being blatantly wrong. Training data isn't garbage. It's programs that do their job, but that are sprinkled with errors. Uncorrelated errors gets averaged out during autoregressive pretraining. Correlated errors can be somewhat suppressed during post-training. Hallucinations (of the generalization-error kind) can be dealt with using synthetic data that improves the model's generalization.