Introduction
Artificial intelligence (AI) characterization (CADx) systems for colonic polyps exhibit high diagnostic accuracy. However, they may only support, not replace, clinical decision-making and little is known of this human-AI interaction – even less so with trainee endoscopists. This issue is difficult to explore, but data from a study investigating CADx-supported trainee endoscopists compared to AI-blinded experts could give insight.
Aims & Methods
This analysis draws on data from a prospective observational study comparing the diagnostic performance of CADx-supported trainee endoscopists (using GI Genius®, Medtronic) to expert endoscopists blinded to AI support. Two key scenarios were explored: (a) instances in which the CADx system produced a ‘no prediction’ output, and (b) cases where trainees actively overruled AI-generated diagnoses. These were assessed to examine the impact of AI absence and human-AI discordance on diagnostic accuracy.
Results
The baseline cohort included 11 trainees performing 225 colonoscopies, mean age 63.8 years (SD 12.7), 48.9% male patients. Baseline ADR was 57.8%. Of 630 resected lesions (median size 3 mm, IQR 2;4), 40.0% were in the rectosigmoid, and 53.6% were non-adenomatous. CADx failed to provide a prediction in 19.2% (n=121) of cases. These ‘no prediction’ polyps were smaller (median 3 mm, p=0.048) and more often located in the right colon (72% vs. 58%, p=0.005. In the absence of AI support, trainee diagnostic accuracy declined significantly by 12.8% (63.6% vs. 76.4%, p=0.007). Interestingly, expert endoscopists—despite being blinded to AI—also performed worse on these polyps (64.7% vs. 73.9%, p=0.045), suggesting an inherent difficulty in classifying these lesions.
Trainees overruled AI suggestion leading to discordant diagnoses in 9.0% (57/630) of polyps, comprised of 20 (35.1%) non-adenomatous, and 37 (64.9%) adenomatous lesions. Trainees incorrectly overruled AI in 59.6% (34/57), which mostly consisted of downgrading lesions to NICE 1. Discordant polyps did not differ in size (3 mm vs. 3 mm, p=0.413), but were more often located proximally (77.2% vs. 58.1%, p=0.005). The difference in false negative rate (FNR) between trainees and AI was 62.16% (81.07% vs. 18.92%, p<0.0001), demonstrating a risk of missed adenomas when trainees override AI. Blinded expert endoscopists achieved identical overall diagnostic accuracy in those polyps as the AI system (59.6%). However, their error profiles differed – In this subset, blinded experts and AI performed equally in overall accuracy (59.6%), although with differing error profiles: AI showed a numerically lower FNR (18.9% vs. 35.1%, Δ16.2%, p=0.17), offset by a higher false positive rate (80.0% vs. 50.0%).
| ´No prediction´ polyps (n=121) | Polyps with AI prediction (n=492) |
|
Size mm, median (IQR)
| 3 (2;4)
| 3 (2;4)
| p=0.0483
|
Location in right colon
| 72%
| 58%
| p=0.005
|
Diagnostic accuracy, trainees, % (95%CI)
| 63.6% (0.55-0.72)
| 76.4% (0.73-0.80)
| p=0.007
|
Diagnostic accuracy, experts, % (95%CI)
| 64.7% (0.53-0.72)
| 73.9% (0.69-0.78)
| p=0.0454
|
Conclusion
CADx systems generate ‘no prediction’ outputs in a substantial proportion of polyps, significantly impairing trainee diagnostic performance. These cases may represent inherently more challenging lesions. Despite CADx’s high accuracy, trainees overrule AI in almost every tenth polyp, most often incorrectly downstaging polyps. Structured training on appropriate use of CADx, emphasizing both its capabilities and limitations, may enhance clinical effectiveness. Further research on human-AI interaction and trust is warranted, as these factors will be critical for successful integration of CADx into routine endoscopic practice.
Disclosure
The original study (AC-CADx) received technical support by Medtronic. Medtronic was not involved in the design of the study or analysis of the data, neither in the original study nor this exploratory analysis.