Introduction
The use of artificial intelligence (AI) systems for decision support has recently gained great popularity in the medicine.1 However, recent studies have shown possible limitations emerging from the interaction between humans and AI. Among these, the term “automation bias (AB)” refers to human errors deriving from over-relying on AI’s false alarms such that the human changes a correct initial interpretation.2 AB is poorly investigated in the medical field and was never tested in the gastroenterological setting.
Aims & Methods
The aim of the study was to investigate the impact of AI on human diagnostic performance in gastrointestinal endoscopy, including AB.
Ninety high-quality endoscopic videoclips (16 seconds) extracted from small bowel (SB) capsule endoscopy (N=60) and proctosigmoidoscopy (N=30) were edited with the addition of a mock AI system designed with diagnostic accuracy, sensitivity and specificity of 80%, 86% and 75%. The ground truth was settled on the agreement of 9 established experts about the presence of SB lesions with significant hemorrhagic potential according to the Saurin classification (30 videoclips), inflammatory lesions relevant for the diagnosis of SB Crohn's disease according to the Lewis score (30 videoclips) and assessing the Ulcerative Colitis (UC) Endoscopic Index of Severity (30 videos). The videoclips were randomly divided into three sessions and submitted around Europe to fully trained endoscopist, trainees and endoscopic technicians. Participants were instructed to the accuracy measures of AI annotations and scored each videoclip first without, then with AI support. Participants’ performances with and without AI were calculated. Subgroup analyses were made after stratification by participants’ experience level in each endoscopic setting and test performances.
Results
A total of 12.660 videoclips (both with and without AI) were visualized in 422 sessions (182, 118 and 122 participants completed the first, second and third session, respectively). A strong and significant difference between the respondents' performance (N=281) when they were unaided by AI (average accuracy (M)=0.78, [0.77, 0.79] and when they were supported by the AI (M=0.81, [0.80, 0.82], paired t-test, t(276)=11.2, p<0.001) was found. Size effect and the evidence in favour of a significant positive effect of AI were large, 0.67 (Cohen's D) and very large (Bayes Factor>100), respectively. Interestingly, no significant difference between trainees and specialists both in regard to their performance, either unaided, M=0.77 [0.74, 0.79] vs M=0.78, [0.77, 0.79] or aided M=0.80 [0.78, 0.82] vs M=0.81 [0.80, 0.82] and their accuracy improvement (M=0.03, [0.02, 0.05] vs M=0.03, [0.02, 0.04] (unaided t-value=-0.89, p=0.37; aided t-value=-0.69, p=0.49, difference t-value=0.54, p=0.59) were found. A strong leveller effect emerged, with low-performers improving their accuracy much more than top-performers (average improvement 5%, [0.04, 0.06] vs 1%, [0.00, 0.02], paired T test, t=6.53, p<0.001). Nine percent of the final decisions involved disregarding correct AI advice in favour of incorrect human judgment, while AB occurred in 2% of cases.
Conclusion
In this study, the application of AI to assess SB bleeding, SB inflammatory lesions, and UC activity brings significant improvement in diagnostic accuracy, with the benefit appearing to be greater in the low-performers. However, in rare cases, AI may induce diagnostic errors due to AB. Future studies directly evaluating the impact of AI in real-life clinical settings are needed to confirm these data.
References
- Cabitza F, Campagner A, Angius R, Natali C, Reverberi C. 2023. AI Shall Have No Dominion: on How to Measure Technology Dominance in AI-supported Human decision-making. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 354, 1–20. https://doi.org/10.1145/3544548.3581095
- Reverberi, C., Rigon, T., Solari, A. et al. Experimental evidence of effective human–AI collaboration in medical decision-making. Sci Rep 12, 14952 (2022). https://doi.org/10.1038/s41598-022-18751-2