Introduction
Endoscopy is essential for the early detection and management of upper gastrointestinal (UGI) tumors, yet variability in diagnostic expertise and workload-related fatigue often compromise the quality of report documentation. These challenges underscore the urgent need for innovative technological approaches to assist endoscopists in optimizing the report generation process.
Multimodal large language models (MLLMs) enable the synergistic integration of visual and semantic representations, demonstrating significant clinical value in medical text generation. This study introduces MERGS, an innovative endoscopy report generation system powered by MLLM. MERGS annotates medical terminology and generates natural language descriptions based on endoscopic images, ultimately producing structured reports that conform to clinical standards.
Aims & Methods
MERGS is built on the pre-trained MLLM Qwen-VL-2.5-7B. It adopts a multimodal architecture integrating visual encoding and cross-modal alignment to enable medical term annotation and natural language generation from endoscopic images. A pathological classification module further enhances lesion characterization. The language model synthesizes image-derived descriptions to generate a comprehensive report with a diagnostic conclusion.
UGI endoscopy reports were collected from the Sun Yat-sen University Cancer Center between 2019 and 2021. A standardized terminology dictionary was developed based on clinical guidelines, supporting a three-tier annotation framework: endoscopic image → standardized medical terminology → natural language text. Model-assisted pre-annotation combined with expert validation substantially reduced manual effort. The resulting dataset was split into training and test sets at an 8:2 ratio.
Model performance was assessed via multi-label classification metrics (accuracy, specificity, F1-score), as well as semantic similarity measures (Jaccard index) to evaluate its ability to recognize anatomical sites and endoscopic features. Natural language generation quality was assessed using BLEU-2 and ROUGE-L scores, and expert satisfaction ratings were collected for 100 randomly sampled outputs to assess readability and clinical relevance. A human–AI comparison was conducted on 50 test reports, comparing system-generated outputs with those authored by experienced endoscopists.
Results
A multimodal dataset was constructed after quality control, comprising 22,373 reports and 135,503 associated images from 21,619 patients. Internal validation demonstrated that MERGS achieved robust performance in automated image interpretation, with 97.3% accuracy and 98.6% specificity for standard anatomical site recognition. In feature identification, it showed high clinical relevance (F1 scores: smooth mucosa 0.915, polyp 0.857, tumor 0.878, erythematous mucosa 0.766).
Generated reports showed strong alignment with reference reports (BLEU-2 = 0.587, ROUGE-L = 0.661) and received high expert satisfaction ratings (mean score: 4.34 ± 0.63 on a 5-point scale). In a human–AI comparison, MERGS-generated reports were considered non-inferior to those authored by endoscopists (scores: MERGS 4.52 ± 0.81 vs. manual 4.68 ± 0.72), with the upper limit of the one-sided 95% CI (0.37) below the non-inferiority margin (0.5, p = 0.0045).
Conclusion
This study developed an automated system for UGI endoscopy reporting powered by an MLLM. Both internal and clinical validation demonstrated the system's high performance in endoscopic image feature recognition and report generation, underscoring its strong potential for clinical deployment.
References
Luo H, Xu G, Li C, He L, Luo L, Wang Z, Jing B, Deng Y, Jin Y, Li Y, Li B, Tan W, He C, Seeruttun S R, Wu Q, Huang J, Huang D W, Chen B, Lin S B, Chen Q M, Yuan C M, Chen H X, Pu H Y, Zhou F, He Y, Xu R H. Real-time artificial intelligence for detection of upper gastrointestinal cancer by endoscopy: a multicentre, case-control, diagnostic study[J]. Lancet Oncol, 2019,20(12):1645-1654.