Abstract
Purpose
Large language models demonstrate increasing utility in healthcare; however, their capacity to handle newly developed medical content beyond training data boundaries remains unclear. This study evaluated the performance of four ChatGPT models on 120 newly developed questions from the 2024 Taiwan Urology Board Examination, all created after the models’ October 2023 training cutoff, thereby reducing the risk of prior model exposure.
Methods
Model performance was assessed using accuracy and processing time. Of 150 examination items, the first 120 newly developed questions constituted the primary analysis set; the remaining 30 previously used archival questions were excluded to reduce potential bias from prior exposure. The primary paired accuracy comparison evaluated o1-preview versus GPT-4o across all 120 questions. Additional analyses across 12 urological subspecialties and question-complexity strata were treated as exploratory.
Results
Across all 120 questions, o1-preview achieved 66.7% accuracy and significantly outperformed GPT-4o (55.8%; p = 0.012), although with longer processing time (19.20 vs. 14.94 s). Performance varied significantly across 12 urological subspecialties (p < 0.001). Unlike the other tested models, o1-preview showed slightly higher accuracy on high-complexity than on low-complexity questions (68.3% vs. 65.0%). On high-complexity surgical anatomy items, o1-preview achieved 80.0% accuracy, whereas GPT-4o and GPT-4o mini both scored 0.0% (p = 0.024).
Conclusion
On this newly developed post-cutoff, text-based urology board question set, o1-preview achieved higher overall accuracy than GPT-4o, particularly on high-complexity surgical anatomy items, at the cost of longer response times. These findings suggest that reasoning-oriented models may offer advantages for selected specialty-level assessment tasks, although generalizability to other languages, modalities, and clinical settings requires further validation.





Similar content being viewed by others
Data availability
The data associated with the paper are available from the corresponding author upon reasonable request.
References
Jeyaraman M, Balaji S, Jeyaraman N, Yadav S (2023) Unraveling the ethical enigma: Artificial intelligence in healthcare. Cureus 15:e43262. https://doi.org/10.7759/cureus.43262
Meskó B, Topol EJ (2023) The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digit Med 6:120. https://doi.org/10.1038/s41746-023-00873-0
Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW (2023) Large language models in medicine. Nat Med 29:1930–1940. https://doi.org/10.1038/s41591-023-02448-8
Hamida SU, Chowdhury MJM, Chakraborty NR, Biswas K, Sami SK (2024) Exploring the landscape of explainable artificial intelligence (XAI): A systematic review of techniques and applications. Big Data Cogn Comput 8:149. https://doi.org/10.3390/bdcc8110149
Brin D, Sorin V, Konen E, Nadkarni G, Glicksberg BS, Klang E (2023) How large language models perform on the united states medical licensing examination: A systematic review. medRxiv. https://doi.org/10.1101/2023.09.03.23294842
Gilson A, Safranek CW, Huang T, Socrates V, Chi L, Taylor RA, Chartash D (2023) How does ChatGPT perform on the United States medical licensing examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 9:e45312. https://doi.org/10.2196/45312
Schoch J, Schmelz H-U, Strauch A, Borgmann H, Nestler T (2024) Performance of ChatGPT-3.5 and ChatGPT-4 on the European board of urology (EBU) exams: A comparative analysis. World J Urol 42:445. https://doi.org/10.1007/s00345-024-05137-4
Tsai CY, Hsieh SJ, Huang HH, Deng JH, Huang YY, Cheng PY (2024) Performance of ChatGPT on the Taiwan urology board examination: Insights into current strengths and shortcomings. World J Urol 42:250. https://doi.org/10.1007/s00345-024-04957-8
Fatima A, Shafique MA, Alam K, Fadlalla Ahmed TK, Mustafa MS (2024) ChatGPT in medicine: A cross-disciplinary systematic review of ChatGPT’s (artificial intelligence) role in research, clinical practice, education, and patient interaction. Medicine 103:e39250. https://doi.org/10.1097/md.0000000000039250
Abuyaman O (2023) Strengths and weaknesses of ChatGPT models for scientific writing about medical vitamin B12: Mixed methods study. JMIR Form Res 7:e49459. https://doi.org/10.2196/49459
Jin M, Yu Q, Shu D, Zhao H, Hua W, Meng Y, Zhang Y, Du M (2024) The impact of reasoning step length on large language models. arXiv. https://doi.org/10.18653/v1/2024.findings-acl.108
Wu S, Peng Z, Du X, Zheng T, Liu M, Wu J, Ma J, Li Y, Yang J, Zhou W Lin Q A comparative study on reasoning patterns of OpenAI’s o1 model? arXiv. https://doi.org/10.48550/arXiv.2410.13639
Temsah MH, Jamal A, Alhasan K, Temsah AA, Malki KH (2024) OpenAI o1-preview vs. ChatGPT in healthcare: A new frontier in medical AI reasoning. Cureus 16:e70640. https://doi.org/10.7759/cureus.70640
Xie Y, Wu J, Tu H, Yang S, Zhao B, Zong Y, Jin Q, Xie C, Zhou Y (2024) A preliminary study of o1 in medicine: Are we closer to an AI doctor? arXiv. https://doi.org/10.48550/arXiv.2409.15277
Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst 35:24824–24837
Soffer S, Sorin V, Nadkarni GN, Klang E (2024) ChatGPT-o1 and the pitfalls of familiar reasoning in medical ethics. https://doi.org/10.1101/2024.09.25.24314342. medRxiv
Liu M, Okuhara T, Dai Z, Huang W, Okada H, Furukawa E, Kiuchi T (2024) Performance of advanced large language models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese medical licensing examination: A comparative study. medRxiv. https://doi.org/10.1101/2024.07.09.24310129
Garg M, Raza S, Rayana S, Liu X, Sohn S (2025) The rise of small language models in healthcare: A comprehensive survey. arXiv. https://doi.org/10.48550/arXiv.2504.17119
Bharatha A, Ojeh N, Fazle Rabbi AM, Campbell M, Krishnamurthy K, Layne-Yarde R, Kumar A, Springer D, Connell K, Majumder MA (2024) Comparing the PERFORMANCE of ChatGPT-4 and medical students on MCQs at varied levels of bloom’s taxonomy. Adv Med Educ Pract Volume 15:393–400. https://doi.org/10.2147/amep.s457408
Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, Liang P (2024) Lost in the middle: How language models use long contexts. Trans Assoc Comput Linguist 12:157–173. https://doi.org/10.1162/tacl_a_00638
Miao J, Thongprayoon C, Suppadungsuk S, Krisanapan P, Radhakrishnan Y, Cheungpasitporn W (2024) Chain of thought utilization in large language models and application in nephrology. Medicina 60:148. https://doi.org/10.3390/medicina60010148
Tsai CW, Lin YJ, Hou JU, Tsai SC, Yeh PC, Kao CH (2025) Optimizing patient education for radioactive iodine therapy and the role of ChatGPT incorporating chain-of-thought technique: ChatGPT questionnaire. Digit Health 11:20552076251357468. https://doi.org/10.1177/20552076251357468
De-Giorgio F, Benedetti B, Mancino M, Sala E, Pascali VL (2025) The need for balancing ’black box’ systems and explainable artificial intelligence: A necessary implementation in radiology. Eur J Radiol 185:112014. https://doi.org/10.1016/j.ejrad.2025.112014
Yudovich MS, Makarova E, Hague CM, Raman JD (2024) Performance of GPT-3.5 and GPT-4 on standardized urology knowledge assessment items in the United States: A descriptive study. J Educ Eval Health Prof 21:17. https://doi.org/10.3352/jeehp.2024.21.17
Kollitsch L, Eredics K, Marszalek M, Rauchenwald M, Brookman-May SD, Burger M, Körner-Riffard K, May M (2024) How does artificial intelligence master urological board examinations? A comparative analysis of different large language models’ accuracy and reliability in the 2022 in-service assessment of the European Board of Urology. World J Urol 42:20. https://doi.org/10.1007/s00345-023-04749-6
Friederichs H, Friederichs WJ, März M (2023) ChatGPT in medical school: How successful is AI in progress testing? Med Educ Online 28:2220920. https://doi.org/10.1080/10872981.2023.2220920
Mogali SR (2024) Initial impressions of ChatGPT for anatomy education. Anat Sci Educ 17:444–447. https://doi.org/10.1002/ase.2261
Amann J, Vayena E, Ormond KE, Frey D, Madai VI, Blasimme A (2023) Expectations and attitudes towards medical artificial intelligence: A qualitative study in the field of stroke. PLoS ONE 18:e0279088. https://doi.org/10.1371/journal.pone.0279088
Sezgin E (2023) Artificial intelligence in healthcare: Complementing, not replacing, doctors and healthcare providers. Digit Health 9:20552076231186520. https://doi.org/10.1177/20552076231186520
Lai VD, Ngo N, Pouran Ben Veyseh A, Man H, Dernoncourt F, Bui T, Nguyen TH (2023) ChatGPT beyond English: Towards a comprehensive evaluation of large language models in multilingual learning. In: Bouamor H, Pino J, Bali K (eds) Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, pp 13171–13189
Acknowledgements
The authors gratefully acknowledge the dedication, collaboration, and intellectual contributions of all co-authors, which were integral to the development and completion of this work.
Funding
This research received no external funding.
Author information
Authors and Affiliations
Contributions
Conceptualization, H. L. and P.-J. C.; methodology, H. L.; software, P.-J. C., H.-K. T. and Y.-C. J.; validation, H. L., C.-L. C. and C.-C. K.; formal analysis, H. L. and M.-H. Y.; investigation, C.-W. T. and E. M.; resources, P.-J. C.; data curation, S.-T. W. and H. L.; writing—original draft preparation, H. L.; writing—review and editing, P.-J. C. and C.-W. T.; visualization, Y.-C. J. and S.-T. W.; supervision, P.-J. C. and M.-H. Y.; project administration, P.-J. C. All authors have read and agreed to the published version of the manuscript.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare no competing interests.
Ethical approval
Not applicable.
Consent to participate
Not applicable.
Consent to publish
Not applicable.
Informed consent
Not applicable.
Additional information
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Information
Below is the link to the electronic supplementary material.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Li, H., Ting, HK., Jhuo, YC. et al. Assessing multiple chatGPT versions on novel content in the Taiwan urology board examination: accuracy, speed, and domain-specific performance. World J Urol 44, 389 (2026). https://doi.org/10.1007/s00345-026-06453-7
Received:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s00345-026-06453-7


