Be careful there. LLMs may be good at identifying a condition based on a description of the symptoms, but they are much worse at recommending the correct course of action (getting it wrong half of the time).
> Participants were randomly assigned to receive assistance from an LLM (GPT-4o, Llama 3, Command R+)
These are pretty old. I'd be curious how performance compares with the latest frontier models.
>GPT-4o, Llama 3, Command R+
The pace of progress is so fast that many studies are totally outdated by the time they release
Going to the doctor for expert advice is often inconvenient, or time consuming, or expensive, or stressful. I think a lot of people want, and seek out, information and advice based on their symptoms, as a first step before a possible doctor's visit. Before LLMs, WebMD (and excessive self-diagnosis based on WebMD) was a meme for a while.
So with regard to LLMs, for me the question is not purely "how often does it get it right?" the question is "how does it compare to the sources and self-diagnosis methods people use otherwise?"
Of course, I agree people should be careful with any form of self-diagnosis or LLM-diagnosis.
Modern models appear to be much better, at least as proxied by their ability to assess urgency in perhaps a more complex setting: mental health (OpenAI benchmark, so perhaps some skepticism is warranted but the methodology seem reasonable and detailed)
I would be careful treating this article as relevant in September 2026. The pace of advancement in the field is such that 6 months is relatively ancient let alone when the models in the study were released. GPT-4o was May 13, 2024. Far before November 2025 when people collectively noticed a turn in LLM value.
LLMs earlier this year were surpassing numerous health related benchmarks when scored against human physicians. Only 2 months after this nature article Harvard posted that LLMs were now outperforming ER docs when given authority to order tests. https://www.harvardmagazine.com/ai/ai-outperforms-doctors-di...
...I know, which is why "as a Language Model and not a real doctor" is a pointless comment to start off with. It should simply not recommend treatment if it's not sure it is correct. I wouldn't blame it or anyone if they asked for help treating a stiff neck, and the LLM (or your neighbor or parent or spouse) suggested light exercises to help relieve it - and do not jump to the suspicion that you may have meningitis.
As a Human, I do not need to know it is a Language Model.
They recruited people to pretend like they had various issues, and then recorded their interactions with the LLMs.
Someone who's been paid a couple of quid to pretend to have a medical condition can easily miss things, and is unlikely to be anywhere near as invested in drilling down to the correct solution as someone who's really suffering.
Another outcome from the study was that the LLMs could do better with the right people driving them. That's not news.
That's still incredibly valuable, though. I suffered from issues that I'd seen doctors for, undergone an upper endoscopy, adjusted my diet, and taken medication for. An LLM suggested my thyroid was at the root of it. My mom confirmed thyroid issues run in our family and just...never thought to tell me.
Just this year, our cat has been having digestive problems. We got special food for her, which she hates, with the suggestion that she'll need to eat it for the rest of her life. Six vet visits later, Fable 5 suggested two tests that my doctor recommended we didn't get. Both found issues that explain her symptoms, and the vet says she will probably only need a supplement and infrequent two week courses of medicine if she has a flare-up.
All that to say, the helplessness of not knowing what's wrong and the people who could know not really caring enough is something that LLMs do a really great job of mitigating. If you don't have any way to know what's wrong with you or a loved one, or how to find out, you're stuck spending a ton of money (in the US at least) and crossing your fingers that someone gets it right.