This study is pretty bad. The comment (https://news.ycombinator.com/item?id=48970182) on the other link with the direct PDF explains the problem well, which is that nothing here being tested is specific to AI systems.
This study gave people access to an LLM that the researchers knew would give incorrect answers to certain questions, and then quizzed people on those questions, with the option to not respond to a given question if they are unsure about the answer.
This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong.
Obviously that person is both more likely to be willing to respond to the question and is more likely to get it wrong!
There are a lot of things I'm very interested in that are specific to modern LLMs and how they affect learning and confidence (sycophancy, cognitive helplessness, etc.).
This study tested none of those. Its experimental setup is not very different than simply substituting the LLM with a textbook with errors.
Why are textbooks relevant here? Even if you repeated the experiment with a textbook instead of the AI and got the same result, what conclusion would you draw from this? The general conclusion of the study seems to be "giving people access to authoritative-seeming but wrong tools for answering questions outside their area of expertise reduces their ability to say they don't know the answer, even when the answer is wrong". So yeah, don't buy bad textbooks for your employees if you don't want them to give you bad textbook answers - but also don't give them AI for things they don't know, perhaps.
I'll also add that even in these simple experimental conditions, I'd bet that having access to a textbook wouldn't have nearly as much of an effect, for a very simple reason: looking up an answer in a textbook is a lot more work than asking an LLM. So when you don't know and aren't forced to answer, I'd bet it's a lot less likely you'd spend the time to look up the answer in the text book. Even more so if the textbook had "this may contain wrong answers!" printed on the cover, like the AIs do.
You could make the point that it’s no different than the textbook example you gave, but people don’t generally use textbooks like that, while out in the world people do use LLMs like that all the time.
The fear is that we can’t tell when the ai advice is bad on these subjects, and as such probably accept confidently terrible advice.
How often do managers just regurgitate ai advice rather than consulting their experts? How often does a person question an expert because the ai said so?
Naturally, the ai will be right some of the time - but it’s really hard to correct for the times the ai is wrong.
> This study is pretty bad.
The study is OK. The article (and the original headline that came with it) is pretty bad because it claims things that the study doesn't. And I guess it is ironic that the TNW article looks 100% AI-generated.
For those curious, the LLM they provided participants with was Step 3.5 Flash: https://huggingface.co/stepfun-ai/Step-3.5-Flash
+1
I'd wager you get similar results if you gave people a version of Google search that purposely gave you bad results. Like, it's framed as an assistant / lookup tool - is it so surprising that people tend to trust it more? Especially since the participants are likely used to using full-powered models and the researchers give them a purposely gimped one (lol)
People are acting rationally when given AI tools to lookup information, their first consumer use case was as a super-powered Google Search
The implication that makes this study relevant is that an LLM is vastly more likely to have factual errors and possibly wildly hallucinate than a proper textbook. If people act the same with both, that IS the actual problem.
All you have to do is go to the technical Reddits to know this is absolutely true. The AI related ones are even worse.
>This is akin to giving someone a textbook on an obscure subject that has certain factual errors, letting them know they can use that textbook in a quiz on that subject, and then quizzing that person on those facts that the textbook gets wrong.
That strikes me as an incredibly appropriate test because LLM’s are unreliable with factual statements. People need to be able to understand that and not treat them like textbooks which are basically 99.9% accurate (let’s please not bicker over the 99.9%. It’s close enough. A major textbook is safe to treat as accurate, an LLM is not).
"this is akin to givin someone a textbook on an obscure subject that has certain factual errors." My brother all LLMs give factual errors so, no this is not a problem with the study. In your fake experiment you are hypothesizing a 100% factual LLM which does not exist.
"This study tested none of those" So the study is bunk because it didn't test your favorite LLM flaws?
>This is akin to giving someone a textbook
Except LLM isn't a textbook, people know that but believe it nonetheless.
You’re saying that if the LLMs were right, humans would have been correct in trusting the machine for things they didn’t know?
The point is that LLMs aren’t right, and the people who took the test were probably reminded of that.
Would people have trusted the textbook you’re mentioning if there was a big red warning on each page that said “this book may contain errors”.
The willingness to trust AI even though it may be wrong and even though there’s money on the line is intersting enough as a study imo
Agreed, the headline says "AI advice made people three times less accurate". But if we really want to know how accurate these people were, we need to know how accurate the AI system they use is. If the AI system is hobbled to a point where it is worse than they reasonably expect we can't blame the people or the AI system. This would be the same as claiming the listening to experts make people less accurate in a study that told experts to lie.
A better headline would read, "Very inaccurate AI made people less accurate" but this would make people naturally ask, "what about a reasonably accurate AI?".