Introduction
LLMs (large language models) already clear recall-focused tests, yet their capacity for nuanced clinical reasoning is still in doubt (1,2). Recently, two innovations have sought to close this gap: reasoning-optimized models that reveal a step-by-step “chain of thought,” and deep-search variants that iteratively comb vast text corpora to build evidence-based answers. Whether explicit reasoning, rather than sheer data volume, actually boosts performance on demanding assessments such as the EBGH (European Board of Gastroenterology and Hepatology) examination remains to be demonstrated.
Aims & Methods
We aimed to evaluate whether reasoning capabilities represent a significant advancement beyond increasing literature exposure during training.For this purpose, we administered 203 EBGH multiple-choice questions to five LLMs: non-reasoning models (ChatGPT-4o, Gemini-1.5-Pro, Llama-3.1-405B, DeepSeek-R1) and reasoning-optimized models (ChatGPT-o1, ChatGPT-o1 Pro). To standardize the querying process for each LLM, we employed the following prompt for every single question:"Act as a gastroenterology doctor. Please read the following question and select the single correct answer choice.”. Each item was run three times; accuracy and answer consistency were recorded. PubMed abstract counts, full texts, free full texts, and total papers across ten GI domains served as proxies for literature density.All tests used official API versions current between January 2024 and 15 February 2025.
Results
In December 2024, ChatGPT-4o answered 164 of 203 EBGH questions correctly, while the reasoning-enhanced o1 reached 171, rescuing 16 items that 4o had missed.When the 39 previously incorrect questions were retested in January 2025, 4o solved 22, leaving 17 that remained unsolved; o1 then solved seven of these,DeepResearch variant solved 5, and o1 Pro solved 15,with 6 items uniquely answered by o1 Pro and 2 stumping every model.Qualitative review showed that reasoning LLMs consistently detected tacit clinical cues and offered broader diagnostic or therapeutic reasoning.Total publication count showed the strongest correlation with overall accuracy(r:0.647, R²:0.418, p:0.167),followed by abstract count and full-text count(both r:0.618,R²:0.382,p:0.196), and free full-text availability(r:0.563, R²:0.317, p=0.255).With updated publication volumes, we found statistically significant correlations between publication metrics and non-reasoning model accuracy, with total publication count(r=0.837, R²:0.701, p:0.026), full-text count(r:0.823, R²:0.677, p:0.033), and abstract count(r=0.819,R²=0.670,p:0.035) all showing significant correlations. Reasoning models demonstrated weaker correlations with publication metrics compared to non-reasoning models, with total publication volume correlations of r:0.312(R²:0.097, p:0.566).We observed weak negative correlations between publication metrics and reasoning model consistency, with free full-text availability showing the strongest negative correlation(r:-0.352,R²:0.124, p:0.514). Interestingly, non-reasoning models showed weak positive correlations with publication metrics, with total publication count(r=0.388,R²=0.151,p=0.466) demonstrating the strongest relationship. However, none of these consistency correlations reached statistical significance.
Conclusion
Explicit reasoning, not deep-search over larger corpora, is the key driver of LLM success on challenging GI board questions. Reasoning-optimized models generalize across sparsely studied domains and may provide safer, domain-agnostic decision support for gastroenterologists.
References
1- Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepaño C, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2):e0000198.
2- Lee P, Bubeck S, Petro J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N Engl J Med. 2023;388(13):1233-9.