Introduction
Large language models, such as ChatGPT-4, have significantly enhanced access to Artificial Intelligence (AI) tools. However, their role in clinical practice of gastroenterology remains uncertain1. Accurate classification of esophageal varices is pivotal in managing patients with portal hypertension as it directly influences prognosis and treatment2.
Aims & Methods
This study aimed to assess the potential of ChatGPT-4 as a diagnostic tool to accurately classify esophageal varices, and to compare its diagnostic accuracy and reliability to experienced and trainee endoscopists.
We collected photos taken during esophagogastroduodenoscopy performed at the Endoscopic Units of six tertiary Italian centers. The images were input into blank sheets of ChatGPT-4, which was asked to classify the varices by size, indicate the presence of red markings, and evaluate image quality in terms of insufflation, clarity and lightning. The same set of images and evaluation criteria was shown to six experienced endoscopists (E1, E2, E3, E4, E5, E6), to four junior endoscopists (J1, J2, J3, J4) and to five trainee endoscopists (T1, T2, T3, T4, T5). The evaluations from the experienced endoscopists were used as a Ground Truth. ChatGPT-4 had not undergone any prior task‐specific training for esophageal varices detection and classification.
Results
We collected 327 photos, of which 279 were considered to be of good quality by the experienced endoscopists. In detecting esophageal varices, ChatGPT-4 showed a lower AUC (0.77, 95% CI 0.66-0.88, Se 86.2%, Sp 68.4%) compared to one of the junior endoscopists (J2 AUC 0.92, 95% CI 0.87-0.98, Se 90.4%, Sp 94.4%; p<0.05) and comparable to other endoscopists (p>0.05). ChatGPT-4 showed a lower AUC (0.75, 95% CI 0.70-0.80, Se 67.6%, Sp 81.7%) than 3/4 junior endoscopist and 3/5 trainee endoscopists in distinguishing small (F1) from large (F2-F3) varices (p<0.05). For the classification of F3 varices, ChatGPT-4 achieved a higher AUC (0.71, 95% CI 0.63-0.80, Se 52.8%, Sp 89.7%) compared to one of the trainee endoscopists (T5 AUC 0.58, 95% CI 0.51-0.65; p<0.05) and comparable to that of all endoscopists (p>0.05). In identifying red signs, ChatGPT-4 (AUC 0.58, 95% CI 0.52-0.65, Se 20.5%, Sp 96.2%) performed less effectively than the endoscopists (p<0.05).
Conclusion
ChatGPT-4 achieved diagnostic accuracy comparable to junior and trainee endoscopists in detecting esophageal varices and classifying F3 varices, but it was overperformed in differentiating small (F1) from large varices (F2–F3) and in identifying red signs. While it cannot replace the clinical judgment of an experienced operator, the model shows potential as a supportive tool for training and preliminary triage of endoscopic images. Further refinement and dedicated training on endoscopic sign interpretation are needed to enhance its reliability and applicability in routine clinical practice.
References
1) Klang E, Sourosh A, Nadkarni GN, Sharif K, Lahat A. Evaluating the role of ChatGPT in gastroenterology: a comprehensive systematic review of applications, benefits, and limitations. Therap Adv Gastroenterol. 2023 Dec 25
2) Nguyen GC, Correia AJ, Thuluvath PJ. The impact of cirrhosis and portal hypertension on mortality following colorectal surgery: a nationwide, population- based study. Dis Colon Rectum. 2009 Aug