Introduction
Deep learning systems in gastrointestinal endoscopy are traditionally pretrained on large, diverse datasets comprising generic, non-medical images before being trained on scarce endoscopic data. Foundation models are models pretrained based on domain-specific datasets. For endoscopy, foundation models require large datasets of general endoscopic images. Yet datasets for developing such models remain limited. In this study, we present GastroNet-5M, a dataset comprising 4,820,653 endoscopic images of approximately 500,000 procedures.
Aims & Methods
GastroNet-5M is composed of anonymized endoscopic images collected from eight Dutch hospitals between 2012 and 2020. Using a self-supervised learning approach, this dataset enabled the development of a foundation model for various endoscopic AI applications. In this study, we compared GastroNet-5M-pretrained models against current state-of-the-art models with traditional pretraining approaches and all publicly available endoscopic foundation models.
We evaluated diagnostic performance on 9 classification and 9 segmentation tasks, including Barrett’s neoplasia detection, colorectal polyp characterization, and gastric cancer invasion depth prediction. All models were trained from scratch for each task. Data efficiency was assessed by repeating all classification tasks with progressively smaller training sets (i.e. 100%, 50%, 25%, 10% and 5%). Robustness was evaluated across 12 datasets reflecting data heterogeneity such as variation in endoscope manufacturer, imaging modality (e.g., virtual chromoendoscopy), and image quality.
Results
GastroNet-5M-pretrained models outperformed all benchmarks in 8 of 9 classification tasks, with an average AUC improvement of 6.8% (Table 1). For segmentation, they achieved higher Dice scores compared to all other models across all 9 tasks, with an average gain of 15.4%. In data efficiency experiments, GastroNet-5M models achieved the highest AUC in 5 of 9 tasks when training data was limited, while no other model led in more than one. In robustness testing, they outperformed other models in 8 of 12 tasks evaluating heterogeneous data conditions.
Table 1. Area-Under-the-Curve scores of the GastroNet-5M-pretrained model and benchmark models for all classification tasks.
Dataset
| Application
| GastroNet-5M FM
| Sanderson et al.1
| Van der Putten et al.2
| EndoFM3 | ConvNeXt Base4
| DINOv2 ViT-Base (ImageNet)5,6
| YOLOv117
|
|---|
| ARGOS8 | Barrett's Neoplasia Detection | 91.6 (87.1-96.1) | 89.9 (86.4-93.4) | 72.6 (69.4-75.9) | 84.9 (83.7-86.1) | 89.2 (87.3-91.0) | 82.2 (77.6-86.8) | 85.2 (83.9-86.5) |
| BONSAI CADe9 | Barrett's Neoplasia Detection | 96.7 (96.6-96.9) | 90.5 (90.2-90.7) | 86.1 (85.0-87.2) | 90.4 (90.0-90.9) | 96.7 (96.5-96.9) | 96.5 (96.3-96.7) | 91.8 (90.7-92.9) |
| BONSAI CADx10 | Barrett's Neoplasia Characterization | 97.2 (96.9-97.6) | 92.1 (91.7-92.4) | 93.1 (92.4-93.9) | 94.4 (94.0-94.7) | 96.4 (96.0-96.8) | 96.8 (96.1-97.5) | 96.2 (95.7-96.6) |
| GIANA11 | Angiodysplasia Detection | 99.5 (99.3-99.6) | 91.7 (91.0-92.3) | 94.7 (94.0-95.4) | 98.7 (98.7-98.8) | 99.6 (99.5-99.7) | 99.6 (99.6-99.7) | 99.4 (99.2-99.6) |
| HyperKvasir-MES12 | IBD Disease Activity | 88.9 (88.7-89.1) | 65.2 (61.8-68.7) | 79.0 (77.4-80.6) | 85.0 (84.2-85.8) | 84.5 (83.9-85.2) | 70.8 (61.2-80.4) | 83.5 (81.8-85.2) |
| LIMUC13 | IBD Disease Activity | 95.5 (95.2-95.7) | 92.7 (92.6-92.8) | 92.1 (91.6-92.7) | 94.2 (94.1-94.4) | 94.3 (94.2-94.5) | 94.5 (94.3-94.8) | 93.3 (92.7-93.8) |
| Maastricht14 | Colorectal Polyp Diagnosis | 92.8 (91.4-94.3) | 73.6 (70.8-76.5) | 76.9 (73.5-80.3) | 90.5 (90.0-91.1) | 89.9 (89.2-90.7) | 81.0 (69.8-92.3) | 89.3 (88.1-90.5) |
| NJDTH15 | Gastric Cancer Invasion Depth | 95.0 (94.1-95.8) | 83.3 (82.3-84.3) | 79.6 (78.3-80.9) | 88.3 (87.8-88.8) | 89.4 (87.8-90.9) | 91.0 (88.4-93.6) | 90.5 (89.2-91.8) |
| POLAR16 | Colorectal Polyp Diagnosis | 85.2 (84.3-86.1) | 67.6 (67.2-68.1) | 64.0 (62.2-65.9) | 71.0 (70.2-71.8) | 77.5 (75.9-79.1) | 67.4 (59.3-75.5) | 77.2 (75.2-79.3) |
Conclusion
GastroNet-5M, a multicenter dataset of over 5 million unlabeled endoscopic images, offers a valuable resource for pretraining deep learning models in endoscopy. The use of GastroNet-5M enhances diagnostic accuracy, reduces the need for scarce application-specific endoscopic imagery, and improves robustness of endoscopic AI systems to the inevitable data heterogeneity in clinical practice. GastroNet-5M is publicly accessible for further research and development.
References
- Sanderson E, Matuszewski BJ. arXiv preprint arXiv:2401.06278. 2024.
- van der Putten J, de Groof J, Struyvenberg M, et al. Artif Intell Med. 2020;107.
- Wang Z, Liu C, Zhang S, et al. arXiv preprint arXiv:2306.16741. 2023.
- Liu Z, Mao H, Wu C-Y, et al. arXiv preprint arXiv:2201.03545. 2022.
- Dosovitskiy A, Beyer L, Kolesnikov A, et al. arXiv preprint arXiv:2010.11929. 2020.
- Oquab M, Darcet T, Moutakanni T, et al. arXiv preprint arXiv:2304.07193. 2023.
- Jocher G, Qiu J. Ultralytics YOLOv11, v11.0.0. 2024.
- de Groof AJ, Struyvenberg MR, van der Putten J, et al. Gastroenterology. 2020;158(4):915–29.e4.
- Fockens KN, Jukema JB, Boers T, et al. United European Gastroenterol J. 2023;11(4):324–36.
- Jukema J, Kusters C, Jong M, et al. Endoscopy (ESGE Days). 2023.
- GIANA Challenge. Available from: https://giana.grand-challenge.org
- Borgli H, Thambawita V, Smedsrud PH, et al. Sci Data. 2020;7(1):283.
- Polat G, Kani HT, Ergenc I, et al. Inflamm Bowel Dis. 2023;29(9):1431–9.
- van der Zander QEW, Schreuder RM, Fonollà R, et al. Endoscopy. 2021.
- Tang D, Zhou J, Wang L, et al. Front Oncol. 2021;11:654093.
- Houwen BB, Hazewinkel Y, Giotis I, et al. Endoscopy. 2022;54(Suppl 1):eP133.