logoalt Hacker News

appplicationtoday at 1:34 PM1 replyview on HN

> Peer review is not perfect, and may not be tuned to catch LLM’s style of errors

This summarizes I think a lot of the challenges with validating LLM output. We hear “humans make mistakes too”, but I would agree with you that our human detection of human-made mistakes and LLM-made mistakes is unlikely to have the same coverage.


Replies

airstriketoday at 2:55 PM

The real problem is that human mistakes generally occur more frequently given the difficulty of the task, whereas LLM mistakes are somewhat random, like the carwash problem, because LLMs cannot truly reason.

I'd rather have a human on my team for whom I can reasonably surmise what tasks they're good at than have a robot who randomly gets shit wrong.