Introduction
Interest in large language models (LLMs) across multiple medical specialities has grown rapidly as a source of patient‑facing advice., No systematic review has appraised the accuracy and readability of LLM‑generated responses for gastrointestinal (GI) procedures, resulting in a lack of current evidence on the reliability of these tools.
Aims & Methods
We aimed to quantify the accuracy and readability of LLM responses to both patient‑ and clinician‑facing questions on GI investigations and interventions, and to identify limitations hindering clinical use. We searched PubMed, Ovid MEDLINE, Ovid Embase, and Google Scholar for peer‑reviewed studies published from 2023 to 2025 that evaluated text-based answers from any LLM on at least one GI diagnostic or therapeutic procedure (colonoscopy, endoscopy, EGD, ERCP, EUS, PEG, GI surgery and others). Inclusion criteria were English language, full‑text primary studies reporting at least one accuracy or readability metric; exclusion criteria included studies with a non-GI focus, image-only tasks, reviews, abstracts, or missing data.. Multiple reviewers independently screened records and extracted data, resolving disagreements by consensus. Risk of bias was appraised using the NIH Quality Assessment Tool (NIH‑QAT) for observational cross‑sectional studies [1]. Accuracy was quantified using a binary (accurate/inaccurate), three-point (accurate, partially accurate, inaccurate), or numerical 1-5 point scales. Readability was assessed using a five‑point Likert scale.
Results
Of the 409 records screened, 45 full‑text articles were reviewed, and 14 met the inclusion criteria, covering seven LLMs. Accuracy varied across models: on the binary scale, Google Bard performed best (82 %, n = 11), followed by Microsoft Copilot (60 %, n = 10). On the three‑point scale, Claude performed best (2.44 / 3, n = 124), with ChatGPT second (2.15 / 3, n = 124). On the five‑point scale, ChatGPT ranked first (3.6 / 5, n = 308), and Claude next (3.04 / 5, n = 16).
By procedure, colonoscopy scored highest on binary (73 %, n = 30) and three‑point (2.73 / 3, n = 14) scales, whereas endoscopy topped the five‑point scale (4.01 / 5, n = 32). Surgery showed mixed performance — binary 59 % (n = 141), three‑point 1.93 / 3 (n = 105), and five‑point 3.54 / 5 (n = 244). ERCP, EUS, PEG, and EGD were represented by smaller samples, giving variable three‑point scores (ERCP 2.66, EUS 2.56, EGD 2.22, PEG 1.83).
| Accuracy | Readability |
Engine | Binary | Three-Point | Five-Point | Five-Point |
ChatGPT | 0.63 (n=140) | 2.15 (n=124) | 3.6 (n=308) | 3.5 (n=308) |
Claude | N/A | 1.61 (n=124) | 3.04 (n=16) | 3.64 (n=16) |
Google Bard | 0.82 (n=11) | 2.44 (n=124) | 2.01 (n=16) | 3.18 (n=16) |
Llama | N/A | 1.64 (n=124) | N/A | N/A |
Microsoft CoPilot | 0.60 (n=10) | N/A | 2.26 (n=16) | 3.26 (n=16) |
Mixtral | N/A | 1.72 (n=124) | N/A | N/A |
Perplexity | 0.20 (n=10) | N/A | N/A | N/A |
Conclusion
LLMs can provide readable answers on GI procedures, but none reach the ≥ 95 % accuracy typically expected for clinical use, with performance still inconsistent across procedures. Standardised benchmarking, balanced prompt libraries, and updated model evaluations are needed before LLMs can be safely used in clinical gastroenterology.
References
[1] National Heart, Lung, and Blood Institute. Study Quality Assessment Tools [https://www. nhlbi. nih. gov/health-topics/study-quality-assessment-tools].