Skip to main content
Log in

Assessing multiple chatGPT versions on novel content in the Taiwan urology board examination: accuracy, speed, and domain-specific performance

  • Research
  • Published:
World Journal of Urology Aims and scope Submit manuscript

Abstract

Purpose

Large language models demonstrate increasing utility in healthcare; however, their capacity to handle newly developed medical content beyond training data boundaries remains unclear. This study evaluated the performance of four ChatGPT models on 120 newly developed questions from the 2024 Taiwan Urology Board Examination, all created after the models’ October 2023 training cutoff, thereby reducing the risk of prior model exposure.

Methods

Model performance was assessed using accuracy and processing time. Of 150 examination items, the first 120 newly developed questions constituted the primary analysis set; the remaining 30 previously used archival questions were excluded to reduce potential bias from prior exposure. The primary paired accuracy comparison evaluated o1-preview versus GPT-4o across all 120 questions. Additional analyses across 12 urological subspecialties and question-complexity strata were treated as exploratory.

Results

Across all 120 questions, o1-preview achieved 66.7% accuracy and significantly outperformed GPT-4o (55.8%; p = 0.012), although with longer processing time (19.20 vs. 14.94 s). Performance varied significantly across 12 urological subspecialties (p < 0.001). Unlike the other tested models, o1-preview showed slightly higher accuracy on high-complexity than on low-complexity questions (68.3% vs. 65.0%). On high-complexity surgical anatomy items, o1-preview achieved 80.0% accuracy, whereas GPT-4o and GPT-4o mini both scored 0.0% (p = 0.024).

Conclusion

On this newly developed post-cutoff, text-based urology board question set, o1-preview achieved higher overall accuracy than GPT-4o, particularly on high-complexity surgical anatomy items, at the cost of longer response times. These findings suggest that reasoning-oriented models may offer advantages for selected specialty-level assessment tasks, although generalizability to other languages, modalities, and clinical settings requires further validation.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Subscribe and save

Springer+
from $39.99 /Month
  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Fig. 1
Fig. 2
Fig. 3
Fig. 4
Fig. 5

Similar content being viewed by others

Data availability

The data associated with the paper are available from the corresponding author upon reasonable request.

References

  1. Jeyaraman M, Balaji S, Jeyaraman N, Yadav S (2023) Unraveling the ethical enigma: Artificial intelligence in healthcare. Cureus 15:e43262. https://doi.org/10.7759/cureus.43262

    Article  PubMed  PubMed Central  Google Scholar 

  2. Meskó B, Topol EJ (2023) The imperative for regulatory oversight of large language models (or generative AI) in healthcare. npj Digit Med 6:120. https://doi.org/10.1038/s41746-023-00873-0

    Article  PubMed  PubMed Central  Google Scholar 

  3. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW (2023) Large language models in medicine. Nat Med 29:1930–1940. https://doi.org/10.1038/s41591-023-02448-8

    Article  CAS  PubMed  Google Scholar 

  4. Hamida SU, Chowdhury MJM, Chakraborty NR, Biswas K, Sami SK (2024) Exploring the landscape of explainable artificial intelligence (XAI): A systematic review of techniques and applications. Big Data Cogn Comput 8:149. https://doi.org/10.3390/bdcc8110149

    Article  Google Scholar 

  5. Brin D, Sorin V, Konen E, Nadkarni G, Glicksberg BS, Klang E (2023) How large language models perform on the united states medical licensing examination: A systematic review. medRxiv. https://doi.org/10.1101/2023.09.03.23294842

    Article  Google Scholar 

  6. Gilson A, Safranek CW, Huang T, Socrates V, Chi L, Taylor RA, Chartash D (2023) How does ChatGPT perform on the United States medical licensing examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med Educ 9:e45312. https://doi.org/10.2196/45312

    Article  PubMed  PubMed Central  Google Scholar 

  7. Schoch J, Schmelz H-U, Strauch A, Borgmann H, Nestler T (2024) Performance of ChatGPT-3.5 and ChatGPT-4 on the European board of urology (EBU) exams: A comparative analysis. World J Urol 42:445. https://doi.org/10.1007/s00345-024-05137-4

    Article  PubMed  Google Scholar 

  8. Tsai CY, Hsieh SJ, Huang HH, Deng JH, Huang YY, Cheng PY (2024) Performance of ChatGPT on the Taiwan urology board examination: Insights into current strengths and shortcomings. World J Urol 42:250. https://doi.org/10.1007/s00345-024-04957-8

    Article  PubMed  Google Scholar 

  9. Fatima A, Shafique MA, Alam K, Fadlalla Ahmed TK, Mustafa MS (2024) ChatGPT in medicine: A cross-disciplinary systematic review of ChatGPT’s (artificial intelligence) role in research, clinical practice, education, and patient interaction. Medicine 103:e39250. https://doi.org/10.1097/md.0000000000039250

    Article  PubMed  PubMed Central  Google Scholar 

  10. Abuyaman O (2023) Strengths and weaknesses of ChatGPT models for scientific writing about medical vitamin B12: Mixed methods study. JMIR Form Res 7:e49459. https://doi.org/10.2196/49459

    Article  PubMed  PubMed Central  Google Scholar 

  11. Jin M, Yu Q, Shu D, Zhao H, Hua W, Meng Y, Zhang Y, Du M (2024) The impact of reasoning step length on large language models. arXiv. https://doi.org/10.18653/v1/2024.findings-acl.108

  12. Wu S, Peng Z, Du X, Zheng T, Liu M, Wu J, Ma J, Li Y, Yang J, Zhou W Lin Q A comparative study on reasoning patterns of OpenAI’s o1 model? arXiv. https://doi.org/10.48550/arXiv.2410.13639

  13. Temsah MH, Jamal A, Alhasan K, Temsah AA, Malki KH (2024) OpenAI o1-preview vs. ChatGPT in healthcare: A new frontier in medical AI reasoning. Cureus 16:e70640. https://doi.org/10.7759/cureus.70640

    Article  PubMed  PubMed Central  Google Scholar 

  14. Xie Y, Wu J, Tu H, Yang S, Zhao B, Zong Y, Jin Q, Xie C, Zhou Y (2024) A preliminary study of o1 in medicine: Are we closer to an AI doctor? arXiv. https://doi.org/10.48550/arXiv.2409.15277

  15. Wei J, Wang X, Schuurmans D, Bosma M, Xia F, Chi E, Le QV, Zhou D (2022) Chain-of-thought prompting elicits reasoning in large language models. Adv Neural Inf Process Syst 35:24824–24837

    Google Scholar 

  16. Soffer S, Sorin V, Nadkarni GN, Klang E (2024) ChatGPT-o1 and the pitfalls of familiar reasoning in medical ethics. https://doi.org/10.1101/2024.09.25.24314342. medRxiv

  17. Liu M, Okuhara T, Dai Z, Huang W, Okada H, Furukawa E, Kiuchi T (2024) Performance of advanced large language models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese medical licensing examination: A comparative study. medRxiv. https://doi.org/10.1101/2024.07.09.24310129

    Article  PubMed  PubMed Central  Google Scholar 

  18. Garg M, Raza S, Rayana S, Liu X, Sohn S (2025) The rise of small language models in healthcare: A comprehensive survey. arXiv. https://doi.org/10.48550/arXiv.2504.17119

    Article  Google Scholar 

  19. Bharatha A, Ojeh N, Fazle Rabbi AM, Campbell M, Krishnamurthy K, Layne-Yarde R, Kumar A, Springer D, Connell K, Majumder MA (2024) Comparing the PERFORMANCE of ChatGPT-4 and medical students on MCQs at varied levels of bloom’s taxonomy. Adv Med Educ Pract Volume 15:393–400. https://doi.org/10.2147/amep.s457408

    Article  Google Scholar 

  20. Liu NF, Lin K, Hewitt J, Paranjape A, Bevilacqua M, Petroni F, Liang P (2024) Lost in the middle: How language models use long contexts. Trans Assoc Comput Linguist 12:157–173. https://doi.org/10.1162/tacl_a_00638

    Article  Google Scholar 

  21. Miao J, Thongprayoon C, Suppadungsuk S, Krisanapan P, Radhakrishnan Y, Cheungpasitporn W (2024) Chain of thought utilization in large language models and application in nephrology. Medicina 60:148. https://doi.org/10.3390/medicina60010148

    Article  PubMed  PubMed Central  Google Scholar 

  22. Tsai CW, Lin YJ, Hou JU, Tsai SC, Yeh PC, Kao CH (2025) Optimizing patient education for radioactive iodine therapy and the role of ChatGPT incorporating chain-of-thought technique: ChatGPT questionnaire. Digit Health 11:20552076251357468. https://doi.org/10.1177/20552076251357468

    Article  PubMed  PubMed Central  Google Scholar 

  23. De-Giorgio F, Benedetti B, Mancino M, Sala E, Pascali VL (2025) The need for balancing ’black box’ systems and explainable artificial intelligence: A necessary implementation in radiology. Eur J Radiol 185:112014. https://doi.org/10.1016/j.ejrad.2025.112014

    Article  PubMed  Google Scholar 

  24. Yudovich MS, Makarova E, Hague CM, Raman JD (2024) Performance of GPT-3.5 and GPT-4 on standardized urology knowledge assessment items in the United States: A descriptive study. J Educ Eval Health Prof 21:17. https://doi.org/10.3352/jeehp.2024.21.17

    Article  PubMed  PubMed Central  Google Scholar 

  25. Kollitsch L, Eredics K, Marszalek M, Rauchenwald M, Brookman-May SD, Burger M, Körner-Riffard K, May M (2024) How does artificial intelligence master urological board examinations? A comparative analysis of different large language models’ accuracy and reliability in the 2022 in-service assessment of the European Board of Urology. World J Urol 42:20. https://doi.org/10.1007/s00345-023-04749-6

    Article  PubMed  Google Scholar 

  26. Friederichs H, Friederichs WJ, März M (2023) ChatGPT in medical school: How successful is AI in progress testing? Med Educ Online 28:2220920. https://doi.org/10.1080/10872981.2023.2220920

    Article  PubMed  PubMed Central  Google Scholar 

  27. Mogali SR (2024) Initial impressions of ChatGPT for anatomy education. Anat Sci Educ 17:444–447. https://doi.org/10.1002/ase.2261

    Article  PubMed  Google Scholar 

  28. Amann J, Vayena E, Ormond KE, Frey D, Madai VI, Blasimme A (2023) Expectations and attitudes towards medical artificial intelligence: A qualitative study in the field of stroke. PLoS ONE 18:e0279088. https://doi.org/10.1371/journal.pone.0279088

    Article  CAS  PubMed  PubMed Central  Google Scholar 

  29. Sezgin E (2023) Artificial intelligence in healthcare: Complementing, not replacing, doctors and healthcare providers. Digit Health 9:20552076231186520. https://doi.org/10.1177/20552076231186520

    Article  PubMed  PubMed Central  Google Scholar 

  30. Lai VD, Ngo N, Pouran Ben Veyseh A, Man H, Dernoncourt F, Bui T, Nguyen TH (2023) ChatGPT beyond English: Towards a comprehensive evaluation of large language models in multilingual learning. In: Bouamor H, Pino J, Bali K (eds) Findings of the Association for Computational Linguistics: EMNLP 2023. Association for Computational Linguistics, Singapore, pp 13171–13189

    Chapter  Google Scholar 

Download references

Acknowledgements

The authors gratefully acknowledge the dedication, collaboration, and intellectual contributions of all co-authors, which were integral to the development and completion of this work.

Funding

This research received no external funding.

Author information

Authors and Affiliations

Authors

Contributions

Conceptualization, H. L. and P.-J. C.; methodology, H. L.; software, P.-J. C., H.-K. T. and Y.-C. J.; validation, H. L., C.-L. C. and C.-C. K.; formal analysis, H. L. and M.-H. Y.; investigation, C.-W. T. and E. M.; resources, P.-J. C.; data curation, S.-T. W. and H. L.; writing—original draft preparation, H. L.; writing—review and editing, P.-J. C. and C.-W. T.; visualization, Y.-C. J. and S.-T. W.; supervision, P.-J. C. and M.-H. Y.; project administration, P.-J. C. All authors have read and agreed to the published version of the manuscript.

Corresponding author

Correspondence to Pei-Jhang Chiang.

Ethics declarations

Conflict of interest

The authors declare no competing interests.

Ethical approval

Not applicable.

Consent to participate

Not applicable.

Consent to publish

Not applicable.

Informed consent

Not applicable.

Additional information

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Supplementary Information

Below is the link to the electronic supplementary material.

Supplementary Material 1 (download ZIP )

Rights and permissions

Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.

Reprints and permissions

About this article

Check for updates. Verify currency and authenticity via CrossMark

Cite this article

Li, H., Ting, HK., Jhuo, YC. et al. Assessing multiple chatGPT versions on novel content in the Taiwan urology board examination: accuracy, speed, and domain-specific performance. World J Urol 44, 389 (2026). https://doi.org/10.1007/s00345-026-06453-7

Download citation

  • Received:

  • Accepted:

  • Published:

  • Version of record:

  • DOI: https://doi.org/10.1007/s00345-026-06453-7

Keywords

Profiles

  1. Chien-Chang Kao