MIT researchers and collaborators found that AI explainability tools in the health sector can produce sharply different results depending on who uses them.
When applied to skin disease diagnosis, non-experts improved their accuracy with AI assistance, although the improvement largely came from deferring to the model. Primary care providers showed a different pattern: they performed best when they received an AI prediction without an explanation.
The study – which appears in Nature Medicine – examined dermatological diagnosis, where AI tools already support some clinicians and increasingly reach patients through AI-powered search products.
Marzyeh Ghassemi, an associate professor in MIT’s Department of Electrical Engineering and Computer Science, said the findings require care in the design of health AI interfaces.
“Good AI systems can improve performance in some health settings, but this has to be balanced carefully with algorithmic deference that can lead to more error,” she said. “We know that both AI and explainability methods can engage automation bias in humans, and this anchoring effect is something that must be accounted for when we design AI systems.”
The interface changes the diagnosis
Explainable AI aims to give users grounds to assess a model’s output. A system may highlight areas of a medical image that influenced its diagnosis. Another approach can show similar images that support a prediction.
Large language models offer a different route. They can produce a plain-language account of a model’s reasoning, presenting a diagnosis in terms intended for a general audience.
The MIT-led research tested several of these approaches. Participants saw medical images alongside an AI prediction of skin disease. One interface supplied a prediction and confidence level without any explanation. Another returned similar images, and a separate system used heat maps to identify regions of interest. Researchers also tested LLM-generated explanations.
Non-experts assessed whether images of skin moles showed cancer. Clinicians faced a broader task: they had to provide a differential diagnosis for dermatological disease.
Non-experts deferred most to language explanations
Every explainability approach improved non-expert accuracy in the study. The tools mainly helped participants identify non-cancerous moles.
Researchers also tested a fairness-constrained model intended to address bias against darker skin tones. That model improved accuracy and reduced diagnostic disparities based on skin tone. The performance gain came with a risk. Non-experts relied heavily on the model’s recommendation, and incorrect model output damaged their performance more than correct output improved it.
“The reason non-expert users are better is because they are more reliant on the models. When the model is wrong, it hurts performance more than it helps performance when the model is right. We were just able to train very good AI models for this setting,” Ghassemi said.
LLM explanations produced the strongest deference effect. Participants trusted those explanations whether the model output was correct or incorrect. They also found vague or generic explanations more convincing, according to the researchers.
Users who received LLM assistance reported greater confidence in wrong answers. That result puts pressure on interface design for consumer-facing diagnostic systems, where a plausible textual explanation can look authoritative even when the model has made an error.
Roxana Daneshjou, an assistant professor of biomedical data science and dermatology at Stanford University, said patients with limited medical knowledge face the greatest exposure to incorrect explainable AI output.
“These findings are important as patients increasingly turn to AI to help with their health care,” she said. “Our findings show that those with the least medical knowledge are most likely to be led astray when explainable AI models give an erroneous output.”
Primary care providers used AI differently
Clinicians did not follow incorrect AI explanations in the same way. They remained resilient when the system produced an erroneous recommendation or explanation. Their strongest performance came from a more limited interface where the system gave clinicians the model’s prediction without an accompanying explanation.
LLM explanations produced the smallest accuracy improvement among the tested explainability methods for clinicians. The result does not show that explanations have no role in clinical practice. It shows that an explanation format suited to a patient or novice may not fit a trained user performing differential diagnosis.
Lead author Orson Xu, an assistant professor in Columbia University’s Department of Biomedical Informatics, said: “It really comes down to how each group uses the explanation. A clinician already has a diagnosis in mind and checks the AI against their own training, so a bad explanation gets caught.
“Meanwhile, a non-expert can use that exact same explanation to form an opinion in the first place, so a plausible, confident-sounding rationale can pull them toward the wrong answer. The same tool ends up being an asset for one user and a liability for another.”
The study argues against treating explainability as a standard interface component that works identically for every role. The user’s baseline expertise affects whether an explanation acts as a check on the model or becomes a substitute for independent judgement.
Timing affects automation bias
The researchers also examined when users saw AI assistance. People became more deferential when the system showed an explanation before they had the opportunity to make their own diagnosis. That finding points to a practical design choice: an interface could ask the user for an initial diagnostic hypothesis, and then provide an AI recommendation that surfaces alternative conditions for consideration.
The study found that users who deferred most to AI were also the weakest performers when they completed the task without AI support. These participants may stand to gain from model assistance, although they also face the greatest risk when the model produces incorrect output.
The research compared human and AI performance across different presentations of disease. AI systems outperformed people when symptoms appeared subtly. Humans performed much better when an image contained atypical symptoms or unrelated features.
Clinician tools may need a direct model output that supports review against professional judgement. Patient-facing tools require particular care around LLM explanations, especially where the system presents a confident narrative for an incorrect recommendation.
See also: PRISM2 model uses clinical dialogue to interpret pathology slides
Want to learn more about AI and big data from industry leaders? Check out AI & Big Data Expo taking place in Amsterdam, California, and London. The comprehensive event is part of TechEx and is co-located with other leading technology events including the Cyber Security & Cloud Expo. Click here for more information.
AI News is powered by TechForge Media. Explore other upcoming enterprise technology events and webinars here.
Read the full article here