Research Article

Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination

Volume: 29 Number: 3 September 30, 2026
TR EN

Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination

Abstract

Objective: Although the integration of large language models (LLMs) into dental education is rapidly increasing, their actual performance in domain-specific assessments remains unclear. This study aimed to evaluate and compare the accuracy of four LLMs (ChatGPT-4.0, Gemini Advanced 1.5 Pro, DeepSeek-V3, and Perplexity) on restorative dentistry questions in a dental specialty examination. Materials and methods: A total of 127 multiple-choice questions from the Turkish Dental Specialty Examination (DUS) conducted between 2012 and 2021 were collected and categorized into 19 content areas. Each question was entered into LLMs in Turkish with standardized instructions. Responses were recorded, and their accuracy was determined according to official answer keys. Statistical differences were analyzed using the appropriate tests. Results: ChatGPT-4.0 had the highest accuracy rate (93.65%), followed by Gemini (82.54%), DeepSeek (71.43%), and Perplexity (65.87%). Significant differences were observed between ChatGPT and DeepSeek (p = 0.027) and Perplexity (p = 0.004), but not between ChatGPT and Gemini (p = 0.118). All models showed higher accuracy in theoretical questions but lower performance in clinically oriented areas such as bleaching and cavity preparation. Conclusion: ChatGPT-4.0 demonstrated the highest overall accuracy among the language models evaluated and shows promise as a supportive tool in theoretical dental education. However, its limited performance in clinical domains underlines the need for careful implementation under physician supervision.

Keywords

Supporting Institution

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Ethical Statement

This study did not involve human participants, animal subjects, or clinical data. Therefore, approval from an Ethics Committee was not required. The research was conducted using publicly available data and artificial intelligence analysis methods only. The authors declare that there are no ethical concerns related to this study.

Thanks

The authors would like to thank all colleagues who contributed to the discussions and feedback during the development of this study.

References

  1. 1. Labadze L, Grigolia M, Machaidze L. Role of AI chatbots in education: systematic literature review. Int J Educ Technol High Educ 2023;20:56.
  2. 2. AlZu’bi S, Mughaid A, Quiam F, Hendawi S. Exploring the capabilities and limitations of ChatGPT and alternative large language models. AIA 2024;2(1):28–37.
  3. 3. Kung TH, Cheatham M, Medenilla A, Sillos C, De Leon L, Elepano C, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health 2023;2(2):e0000198.
  4. 4. Hadi MU, Al-Tashi Q, Qureshi R, Shah A, Muneer A, Irfan M, et al. Large language models: a comprehensive survey of applications, challenges, limitations, and future prospects. Authorea Preprints 2025.
  5. 5. ALI K, Barhom N, Marino FT, Duggal M. The thrills and chills of ChatGPT: Implications for assessments in undergraduate dental education. Preprints 2023.
  6. 6. Rao A, Pang M, Kim J, Kamineni M, Lie W, Prasad AK, et al. Assessing the utility of ChatGPT throughout the entire clinical workflow: development and usability study. J Med Internet Res. 2023;25:e48659.
  7. 7. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med 2023;29:1930–1940.
  8. 8. van Dis EAM, Bollen J, Zuidema W, van Rooij R, Bockting CL. ChatGPT: five priorities for research. Nature 2023;614(7947):224–226.

Details

Primary Language

English

Subjects

Restorative Dentistry

Journal Section

Research Article

Publication Date

September 30, 2026

Submission Date

October 1, 2025

Acceptance Date

September 7, 2026

Published in Issue

Year 2026 Volume: 29 Number: 3

APA
Bulut, Ç., Çolak, G., & Çolak, G. (2026). Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination. Cumhuriyet Dental Journal, 29(3), 385-392. https://doi.org/10.7126/cumudj.1794852
AMA
1.Bulut Ç, Çolak G, Çolak G. Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination. Cumhuriyet Dent J. 2026;29(3):385-392. doi:10.7126/cumudj.1794852
Chicago
Bulut, Çilem, Gülben Çolak, and Gürkan Çolak. 2026. “Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination”. Cumhuriyet Dental Journal 29 (3): 385-92. https://doi.org/10.7126/cumudj.1794852.
EndNote
Bulut Ç, Çolak G, Çolak G (September 1, 2026) Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination. Cumhuriyet Dental Journal 29 3 385–392.
IEEE
[1]Ç. Bulut, G. Çolak, and G. Çolak, “Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination”, Cumhuriyet Dent J, vol. 29, no. 3, pp. 385–392, Sept. 2026, doi: 10.7126/cumudj.1794852.
ISNAD
Bulut, Çilem - Çolak, Gülben - Çolak, Gürkan. “Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination”. Cumhuriyet Dental Journal 29/3 (September 1, 2026): 385-392. https://doi.org/10.7126/cumudj.1794852.
JAMA
1.Bulut Ç, Çolak G, Çolak G. Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination. Cumhuriyet Dent J. 2026;29:385–392.
MLA
Bulut, Çilem, et al. “Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination”. Cumhuriyet Dental Journal, vol. 29, no. 3, Sept. 2026, pp. 385-92, doi:10.7126/cumudj.1794852.
Vancouver
1.Çilem Bulut, Gülben Çolak, Gürkan Çolak. Comparative Evaluation of Large Language Models’ Accuracy in Answering Multiple-Choice Restorative Dentistry Questions From a National Specialty Examination. Cumhuriyet Dent J. 2026 Sep. 1;29(3):385-92. doi:10.7126/cumudj.1794852

Cumhuriyet Dental Journal (Cumhuriyet Dent J, CDJ) is the official publication of Cumhuriyet University Faculty of Dentistry. CDJ is an international journal dedicated to the latest advancement of dentistry. The aim of this journal is to provide a platform for scientists and academicians all over the world to promote, share, and discuss various new issues and developments in different areas of dentistry. First issue of the Journal of Cumhuriyet University Faculty of Dentistry was published in 1998. In 2010, journal's name was changed as Cumhuriyet Dental Journal. Journal’s publication language is English.


CDJ accepts articles in English. Submitting a paper to CDJ is free of charges. In addition, CDJ has not have article processing charges.

Frequency: Four times a year (March, June, September, and December)

IMPORTANT NOTICE

All users of Cumhuriyet Dental Journal should visit to their user's home page through the "https://dergipark.org.tr/tr/user" " or "https://dergipark.org.tr/en/user" links to update their incomplete information shown in blue or yellow warnings and update their e-mail addresses and information to the DergiPark system. Otherwise, the e-mails from the journal will not be seen or fall into the SPAM folder. Please fill in all missing part in the relevant field.

Please visit journal's AUTHOR GUIDELINE to see revised policy and submission rules to be held since 2020.