logoalt Hacker News

habineroyesterday at 12:04 AM5 repliesview on HN

If you ask that, you fundamentally misunderstand the point.

It's not about the LLM, it's about whether people will critically evaluate what it spits out.


Replies

adroitbossyesterday at 12:42 AM

If the source is a person instead of an LLM, you still wouldn't be able to evaluate what was said. This is nothing new.

show 2 replies
westoncbyesterday at 12:24 AM

That's fine as a point but it's not what the headline describes. The question of total/real effect on accuracy is also something one could ask about. Both are valid.

encomiastyesterday at 12:39 AM

Here is what the study says:

"The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions."

So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.

The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception.

If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.

show 4 replies
protocoltureyesterday at 1:07 AM

>It's not about the LLM, it's about whether people will critically evaluate what it spits out.

Its about whether people will critically evaluate any information they are given. It has nothing to do with LLMs.

s1artibartfastyesterday at 3:26 AM

How does it test that at all? Did the quiz have answers that people could figure out better by scrutinizing the llm?