Research Article

Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements

Volume: 29 Number: 3 September 30, 2026
EN TR

Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements

Abstract

Objectives: This study aimed to compare the performance of DeepSeek-V3, ChatGPT 4.0, ChatGPT 4.5, Gemini Advanced 2.0, and Grok in answering single-correct answer multiple-choice (SCQ), multiple-correct answer multiple-choice (MCQ), binary (true/false) (BQ), and open-ended questions (OEQ) based on the European Society of Endodontology (ESE) position statements. Materials and Methods: A total of 70 questions (24 SCQ/MCQ, 29 BQ, and 17 OEQ) were posed to five LLMs. Responses were scored on a 0–10 scale for comprehensiveness, scientific accuracy, clarity, and relevance. SCQ, MCQ, and BQ formats were additionally evaluated for correctness (1 = correct, 0 = incorrect). Data were analyzed using Pearson’s chi-square test and the Kruskal-Wallis H test (p < 0.05). Results: In the SCQ/MCQ format, Gemini Advanced 2.0 scored significantly higher than Grok. In the BQ format, ChatGPT 4.5 and Gemini Advanced 2.0 achieved significantly higher scores than ChatGPT 4.0 and Grok. In the OEQ format, ChatGPT 4.5 obtained significantly higher scores than Grok (p < 0.05). Accuracy rates were 67.9% for ChatGPT 4.5, 66.0% for Gemini Advanced 2.0, 58.5% for DeepSeek-V3, 52.8% for ChatGPT 4.0, and 49.1% for Grok (p > 0.05). In all models, scientific accuracy scores were significantly lower than relevance scores (p < 0.05). Conclusions: ChatGPT 4.5 and Gemini Advanced 2.0 demonstrated stronger performance than the other models, whereas Grok showed lower response-content performance. However, since LLMs may appear convincing even when generating incorrect clinical information, they should be used only as supplementary tools in endodontic education and preliminary information retrieval. Clinically relevant outputs should be verified by an endodontist or an appropriately qualified clinician.

Keywords

Ethical Statement

This study was based exclusively on AI-generated responses and publicly available professional position statements. No human participants, patient data, animal subjects, or identifiable personal information were used. Therefore, ethics committee approval and informed consent were not applicable.

Thanks

This study was presented as an oral presentation at the 28th International Dental Congress of the Turkish Dental Association, held in Diyarbakır, Türkiye, on September 18–21, 2025

References

  1. 1. King MR. The future of AI in medicine: a perspective from a chatbot. Ann Biomed Eng 2023;51:291-295.
  2. 2. Dhopte A, Bagde H. Smart smile: revolutionizing dentistry with artificial intelligence. Cureus 2023;15:e41227.
  3. 3. Aminoshariae A, Kulild J, Nagendrababu V. Artificial intelligence in endodontics: current applications and future directions. J Endod 2021;47:1352-1357.
  4. 4. Rodrigues JA, Krois J, Schwendicke F. Demystifying artificial intelligence and deep learning in dentistry. Braz Oral Res 2021;35:e094.
  5. 5. Pompili D, Richa Y, Collins P, Richards H, Hennessey DB. Using artificial intelligence to generate medical literature for urology patients: a comparison of three different large language models. World J Urol 2024;42:455.
  6. 6. Rokhshad R, Zhang P, Mohammad-Rahimi H, Pitchika V, Entezari N, Schwendicke F. Accuracy and consistency of chatbots versus clinicians for answering pediatric dentistry questions: a pilot study. J Dent 2024;144:104938.
  7. 7. Kipp M. From GPT-3.5 to GPT-4o: a leap in AI’s medical exam performance. Information 2024;15:543.
  8. 8. OpenAI. Introducing GPT-4.5. Available from: https://openai.com/index/introducing-gpt-4-5/. Accessed April 1, 2025.

Details

Primary Language

English

Subjects

Endodontics

Journal Section

Research Article

Publication Date

September 30, 2026

Submission Date

May 17, 2026

Acceptance Date

August 31, 2026

Published in Issue

Year 2026 Volume: 29 Number: 3

APA
Defişet, M., Yeniçeri Özata, M., & Çam, A. (2026). Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements. Cumhuriyet Dental Journal, 29(3), 487-496. https://doi.org/10.7126/cumudj.1953387
AMA
1.Defişet M, Yeniçeri Özata M, Çam A. Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements. Cumhuriyet Dent J. 2026;29(3):487-496. doi:10.7126/cumudj.1953387
Chicago
Defişet, Merve, Merve Yeniçeri Özata, and Ayşenur Çam. 2026. “Evaluation of Large Language Model Responses to Different Types of Questions Based on European Society of Endodontology Position Statements”. Cumhuriyet Dental Journal 29 (3): 487-96. https://doi.org/10.7126/cumudj.1953387.
EndNote
Defişet M, Yeniçeri Özata M, Çam A (September 1, 2026) Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements. Cumhuriyet Dental Journal 29 3 487–496.
IEEE
[1]M. Defişet, M. Yeniçeri Özata, and A. Çam, “Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements”, Cumhuriyet Dent J, vol. 29, no. 3, pp. 487–496, Sept. 2026, doi: 10.7126/cumudj.1953387.
ISNAD
Defişet, Merve - Yeniçeri Özata, Merve - Çam, Ayşenur. “Evaluation of Large Language Model Responses to Different Types of Questions Based on European Society of Endodontology Position Statements”. Cumhuriyet Dental Journal 29/3 (September 1, 2026): 487-496. https://doi.org/10.7126/cumudj.1953387.
JAMA
1.Defişet M, Yeniçeri Özata M, Çam A. Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements. Cumhuriyet Dent J. 2026;29:487–496.
MLA
Defişet, Merve, et al. “Evaluation of Large Language Model Responses to Different Types of Questions Based on European Society of Endodontology Position Statements”. Cumhuriyet Dental Journal, vol. 29, no. 3, Sept. 2026, pp. 487-96, doi:10.7126/cumudj.1953387.
Vancouver
1.Merve Defişet, Merve Yeniçeri Özata, Ayşenur Çam. Evaluation of large language model responses to different types of questions based on European Society of Endodontology position statements. Cumhuriyet Dent J. 2026 Sep. 1;29(3):487-96. doi:10.7126/cumudj.1953387

Cumhuriyet Dental Journal (Cumhuriyet Dent J, CDJ) is the official publication of Cumhuriyet University Faculty of Dentistry. CDJ is an international journal dedicated to the latest advancement of dentistry. The aim of this journal is to provide a platform for scientists and academicians all over the world to promote, share, and discuss various new issues and developments in different areas of dentistry. First issue of the Journal of Cumhuriyet University Faculty of Dentistry was published in 1998. In 2010, journal's name was changed as Cumhuriyet Dental Journal. Journal’s publication language is English.


CDJ accepts articles in English. Submitting a paper to CDJ is free of charges. In addition, CDJ has not have article processing charges.

Frequency: Four times a year (March, June, September, and December)

IMPORTANT NOTICE

All users of Cumhuriyet Dental Journal should visit to their user's home page through the "https://dergipark.org.tr/tr/user" " or "https://dergipark.org.tr/en/user" links to update their incomplete information shown in blue or yellow warnings and update their e-mail addresses and information to the DergiPark system. Otherwise, the e-mails from the journal will not be seen or fall into the SPAM folder. Please fill in all missing part in the relevant field.

Please visit journal's AUTHOR GUIDELINE to see revised policy and submission rules to be held since 2020.