Scientific evaluation of the effectiveness of artificial intelligence-based chatbots in prosthodontics
2025
0 views
0 downloads
Advisor: Doç. Dr. Kaan Yerliyurt
Abstract (EN)
Objective: The aim of this study was to scientifically evaluate the effectiveness of six different artificial intelligence-based large language models (LLMs) (ChatGPT-4o, Claude 3.7 Sonnet, Microsoft Copilot, DeepSeek-V2, Gemini 2.0, and Grok-3) in responding to frequently asked patient questions in the field of prosthodontics. The responses were comparatively analyzed based on the parameters of scientific accuracy, comprehensiveness, clarity, and relevance. Materials and Methods: In this in-vitro descriptive and comparative study, 10 standard patient questions frequently encountered in prosthodontic practice were posed to six current LLMs. A total of 60 unique responses were obtained by prompting the models to answer with the persona of a "prosthodontist." The responses were evaluated by two independent academics from the department of prosthodontics using a double-blind method with a predefined 5-point Likert scale (1: Very Poor, 5: Very Good) and a detailed scoring rubric. Statistical analysis of the data was performed to assess inter-rater reliability (Intraclass Correlation Coefficient - ICC), internal consistency of the scale (Cronbach's Alpha), and to compare the performance of the models (Kruskal-Wallis H Test). The significance level was set at p<0.05. Results: The tested LLMs were found to exhibit a moderate-to-good level of performance overall. Claude 3.7 Sonnet achieved the highest overall score (3.79 ± 0.74), while Microsoft Copilot showed the lowest performance (3.29 ± 0.53). While no statistically significant difference was found among the models for the parameters of scientific accuracy (p=0.320), clarity (p=0.184), and relevance (p=0.608), a significant difference was identified in the comprehensiveness parameter (p=0.036). Inter-rater reliability (ICC=0.709, p<0.001) and internal consistency of the scale (Cronbach α=0.709) were found to be "good." The Gemini 2.0 and DeepSeek models demonstrated a high level of positive correlation (r>0.80, p<0.01) among the evaluation parameters. Conclusion: AI-based chatbots have the potential to be valuable auxiliary tools for patient education in the field of prosthodontics. However, there are differences in the depth of the models' responses, and none reached the "very good" category. These technologies cannot replace professional dental consultation in clinical decision-making processes. The choice of model is of critical importance when utilizing LLMs.
Author
Dr. Ahmet Doğan Işık
How to Cite
Ahmet Doğan Işık (Dentistry Specialty Thesis). Scientific evaluation of the effectiveness of artificial intelligence-based chatbots in prosthodontics, 2025, Tokat Gaziosmanpaşa Üniversity.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Tokat Gaziosmanpaşa Üniversity
- Fundamental solutions of a discontinuous conformable boundary value problem(2023)
- COVID-19 hastalarında ACE gen polimorfizminin belirlenmesi(2024)
- Evaluation of the insecticidal effect of some plant extracts and nanoparticles on spodoptera littoralis (Boisd.) (Lepidoptera: Noctuidae) larvae(2024)
- Kelam Bilimi ve zihinsel, psikolojik ve ruhsal yönleri üzerindeki etkileri(2021)
- 2018 Turkish Republic of revolution history course teacher's views on curriculum (Example of Yozgat province)(2019)
- Investigation of the aquaporine molecules expressions in human sperm cells from different age groups(2019)
