Introduction
UC-SCALE[1] is an AI algorithm designed to assess endoscopic severity in ulcerative colitis (UC). It was trained on a large subset of Phase III Etrolizumab (ETRO) clinical trial data from 471 sites. In our previous work, UC-SCALE previously demonstrated excellent agreement (quadratic-weighted kappa >0.80) with centrally read Mayo Clinic Endoscopic Subscores (MCES) on an ETRO test set spanning 83 sites. While clinical trial data is highly controlled, its generalizability to real-world settings remained untested. Here, we assess UC-SCALE on a retrospective cohort from the Paris IBD Center, which we refer to as the Real World Dataset (RWD).
Aims & Methods
The RWD contained 276 unique UC patients (mean 1.17 visits/patient) from the Paris IBD center. For each colonoscopy video, two or three local gastroenterologists reviewed the recordings and reached a consensus MCES (0–3) per colonic segment (descending, sigmoid, rectum). Using the same, unmodified UC-SCALE model trained on ETRO, we evaluated both the RWD and ETRO test sets with the identical assessment protocol detailed in [1]. Quadratic-Weighted Kappa (QWK) quantified the agreement between ordinal predictions made by UC-SCALE and human scoring. Area Under the Receiver Operating Characteristic Curve (AUC) was used to measure the discriminative ability for each MCES class.
Results
On these real-world data, UC-SCALE’s section-level predictions closely mirrored local expert scoring. QWK reached ~0.804 - virtually identical to ETRO (~0.802). For class-specific discrimination (MCES 0–3), AUC values exceeded 0.85 in all classes and in several cases surpassed ETRO’s results. The average AUC across all classes approached 0.90 in the RWD, aligning well with performance in trial conditions. As in ETRO, performance was strongest for Mayo 0 and 3, but notably, predictions for intermediate scores (Mayo 1 and 2) were even more accurate in RWD. These results highlight UC-SCALE’s robust generalizability to routine clinical practice.
Metric | RWD | ETRO | Difference |
Quadratic Weighted Kappa | 0.80 (0.77, 0.84) | 0.80 (0.78, 0.82) | 0.00 |
AUC Mayo 0 | 0.94 (0.93, 0.96) | 0.94 (0.93, 0.95) | -0.00 |
AUC Mayo 1 | 0.85 (0.81, 0.88) | 0.81 (0.78, 0.83) | 0.04 |
AUC Mayo 2 | 0.90 (0.88, 0.92) | 0.82 (0.80, 0.84) | 0.08 |
AUC Mayo 3 | 0.92 (0.89, 0.94) | 0.91 (0.90, 0.93) | 0.01 |
AUC Mean | 0.90 (0.89, 0.92) | 0.87 (0.86, 0.88) | 0.03 |
Table 1: Performance metrics for UC-SCALE on Real-World Data (RWD) versus the Etrolizumab (ETRO) trial dataset. The performance values are presented with 95% confidence intervals, which were derived through bootstrapping using 10,000 resamples.
Conclusion
Applying the original ETRO-trained UC-SCALE model to a real-world UC cohort confirmed its high accuracy and reliability for objective MCES scoring outside of clinical trials. To date, our analysis has focused on predictions at the level of anatomical segments. Future work will expand toward full video-level scoring (closer to clinical workflows) and artificial segments, which offer a more granular topological severity map. Additionally, upcoming analyses will explore correlations between UC-SCALE and key IBD biomarkers, as well as direct comparisons with human scoring.
References
[1] Gutierrez-Becker, Benjamin, et al. "Ulcerative Colitis Severity Classification and Localized Extent (UC-SCALE): An Artificial Intelligence Scoring System for a Spatial Assessment of Disease Severity in Ulcerative Colitis." Journal of Crohn's and Colitis 19.1 (2025): jjae187.
Disclosure
Alvaro Gomariz, Bernhard Stimpel, Benjamin Gutierrez-Becker, Stefan Frässle, Emily Fisher, Steven Levitte and Teresa Karrer are employees of F. Hoffmann-La Roche Ltd.