Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination
Abstract
Objective: Although the integration of large language models (LLMs) into dental education is rapidly increasing, their actual performance in domain-specific assessments remains unclear. This study aimed to evaluate and compare the accuracy of four LLMs (ChatGPT-4.0, Gemini Advanced 1.5 Pro, DeepSeek-V3, and Perplexity) on restorative dentistry questions in a dental specialty examination. Materials and methods: A total of 127 multiple-choice questions from the Turkish Dental Specialty Examination (DUS) conducted between 2012 and 2021 were collected and categorized into 19 content areas. Each question was entered into LLMs in Turkish with standardized instructions. Responses were recorded, and their accuracy was determined according to official answer keys. Statistical differences were analyzed using the appropriate tests. Results: ChatGPT-4.0 had the highest accuracy rate (93.65%), followed by Gemini (82.54%), DeepSeek (71.43%), and Perplexity (65.87%). Significant differences were observed between ChatGPT and DeepSeek (p = 0.027) and Perplexity (p = 0.004), but not between ChatGPT and Gemini (p = 0.118). All models showed higher accuracy in theoretical questions but lower performance in clinically oriented areas such as bleaching and cavity preparation. Conclusion: ChatGPT-4.0 demonstrated the highest overall accuracy among the language models evaluated and shows promise as a supportive tool in theoretical dental education. However, its limited performance in clinical domains underlines the need for careful implementation under physician supervision.
Keywords
Supporting Institution
Ethical Statement
Thanks
References
- 1. Labadze L, Grigolia M, Machaidze L. Role of AI chatbots in education: systematic literature review. Int J Educ Technol High Educ 2023;20:56.
- 2. AlZu’bi S, Mughaid A, Quiam F, Hendawi S. Exploring the capabilities and limitations of ChatGPT and alternative large language models. AIA 2024;2(1):28–37.
- 3. Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepano C, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health 2023;2(2):e0000198.
- 4. Hadi MU, Al-Tashi Q, Qureshi R, Shah A, Muneer A, Irfan M, et al. Large language models: a comprehensive survey of applications, challenges, limitations, and future prospects. Authorea Preprints 2025.
- 5. ALI K, Barhom N, Marino FT, Duggal M. The thrills and chills of ChatGPT: Implications for assessments in undergraduate dental education. Preprints 2023.
- 6. Rao A, Pang M, Kim J, Kamineni M, Lie W, Prasad AK, et al. Assessing the utility of ChatGPT throughout the entire clinical workflow: development and usability study. J Med Internet Res. 2023;25:e48659.
- 7. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med 2023;29:1930–1940.
- 8. van Dis EAM, Bollen J, Zuidema W, van Rooij R, Bockting CL. ChatGPT: five priorities for research. Nature 2023;614(7947):224–226.
Details
Primary Language
English
Subjects
Restorative Dentistry
Journal Section
Research Article
Authors
Çilem Bulut
0000-0001-9436-8399
Türkiye
Gülben Çolak
*
0000-0003-0786-9609
Türkiye
Gürkan Çolak
0009-0000-8393-0913
Sweden
Publication Date
September 30, 2026
Submission Date
October 1, 2025
Acceptance Date
September 7, 2026
Published in Issue
Year 2026 Volume: 29 Number: 3