Following up on this because I've been thinking about why this is. I believe a big part of it is just that the intuitive parts of medicine are learned essentially by internship and therefore are not well encoded into the models. Unlike software, where you can have deterministic output that the models can train on, medical outputs are very fluid and dynamic. What works well in a paper, even though we claim to do evidence-based medicine, may be very far from what a "good" physician does in practice.
The issue is that the AI is at best what's in papers and medical records, which often forgoes the core thing that might lead a physician to uncover something or take a different approach with the patient.