Performance of artificial intelligence in critical patient diagnosis: Comparison of ChatGPT, Gemini, Claude and emergency physicians
2025
0 views
0 downloads
Advisor: Doç. Dr. Sinan Paslı
Abstract (EN)
Introduction: This study aimed to compare the diagnostic process of emergency physicians with the diagnostic outputs generated by ChatGPT, Claude, and Gemini in critically ill patients presenting to the emergency department(ED). Methods: For critically ill patients presenting to the ED, physicians identified five preliminary diagnoses based on vital signs and medical history. The same anonymized data were entered into AI models, which were asked to generate five preliminary diagnoses. After physical examination, laboratory, and imaging results became available, both physicians and AI models refined these into three differential diagnoses. Finally, the physician's final diagnosis, explaining the patient's clinical condition, was accepted as the gold standard. Each AI model was then asked to determine the final diagnosis using the complete dataset. The preliminary, differential, and final diagnoses generated by each model were compared with the gold standard. Statistical analyses were performed using R. Results: The study included 180 patients (56% male, 44% female). Regarding the accuracy of final diagnoses, GPT-4o achieved an accuracy rate of 67.2%, GPT-5 achieved 65.6%, Claude achieved 63.3%, and Gemini achieved 59.4%. No statistically significant difference was found among the groups (p=0.16).When evaluating the agreement between the AI models' definitive diagnoses and the gold-standard diagnosis, Cohen's κ coefficients were calculated as 0.656 for GPT-4o, 0.638 for GPT-5, 0.616 for Claude, and 0.575 for Gemini. Conclusion The finding that AI models achieved diagnostic accuracy rates of 60%-70% and showed overall substantial agreement with the gold standard decision supports their potential to assist clinical decision-making in the ED setting. Keywords: Artificial intelligence, ChatGPT, Claude, Gemini, Large language model
Author
Dr. İbrahim Günaydın
How to Cite
İbrahim Günaydın (Medical Specialty Thesis). Performance of artificial intelligence in critical patient diagnosis: Comparison of ChatGPT, Gemini, Claude and emergency physicians, 2025, Karadeniz Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Karadeniz Technical University
- Optimization of gold recovery from placer deposits using gravity methods(2025)
- Harşit çayından (Tirebolu-Giresun) elde edilen kırılmış dere malzemesinin beton agregası olarak kullanılabilirliğinin incelenmesi(2005)
- Yaşlandırma Süresinin Zn-27Al-1Cu Alaşımının Yapı ve Mekanik Özelliklerine Etkisi(2016)
- Prevalence and associated factors of tobacco use, alcohol consumption, alcohol use disorder among individuals aged 20 and above living in trabzon province(2025)
- Trabzon güney çevre yolu güzergahı Darıca (Akçaabat) - Yalı mahallesi (Trabzon) arasının mühendislik jeolojisi / Investigation of the planned route of the southern highway between Darıca (Akçaabat) - Yalı mahallesi (Trabzon) in terms of engineering geolog(2001)
- EXPERIMENTAL AND NUMERICAL INVESTIGATION OF FATIGUE BEHAVIOUR IN RADIAL JOURNAL BEARINGS MANUFACTURED FROM ZA-27 ALLOY(2025)
