> AI is at least as good at diagnostics as fully qualified doctors
If the prompt is an expert-written board question! Not so with inferior prompts [0]. Critically, you need deep medical knowledge to interact correctly with the agent.
What you're asking is basically: "If we take someone out of a three month dev bootcamp, and have them prompt Claude, why can't they be as good as a four year CS grad?"
I doubt that you would feel similarly about expertise in your own field.
There have been a huge progress in AI reasoning in the past 2 years. GPT-4o mentioned in the article would struggle with high-school math problems, OTOH GPT-6 can solve problems beyond capability of professional mathematicians.
I'd wager GPT-6 would not depend on high-quality prompts, although it might still be good to get a trained person to enter information and do a sanity check.