logoalt Hacker News

slibhblast Sunday at 11:46 PM3 repliesview on HN

> but also don't give them AI for things they don't know, perhaps

The study doesn't show that at all. It didn't test actual AI.

They could have tested a cohort of subjects with access to actual ChatGPT. Ask yourself why they didn't.


Replies

dnemmersyesterday at 11:42 AM

Old AI is so bad it should be disregarded, but new AI is so good, you don't even have to verify its output....

Is that what you're selling us?

So in 18 months, we'll just rinse and repeat?

beepbooptheorylast Sunday at 11:58 PM

Because this is exactly what they controlled for. FTA:

> The researchers used Step 3.5 Flash, a model that was usually wrong on these questions, precisely so any reduction in judgment could not be explained as sensible delegation to a reliable tool.

(emphasis mine)

show 1 reply
wonnageyesterday at 12:01 AM

They provide a sample of hallucinated answers from ChatGPT at the end of the study.