Medical SpecialtyOpen Access

Performance of artificial intelligence in critical patient diagnosis: Comparison of ChatGPT, Gemini, Claude and emergency physicians

2025
0 views
0 downloads
Advisor: Doç. Dr. Sinan Paslı

Abstract (EN)

Introduction: This study aimed to compare the diagnostic process of emergency physicians with the diagnostic outputs generated by ChatGPT, Claude, and Gemini in critically ill patients presenting to the emergency department(ED). Methods: For critically ill patients presenting to the ED, physicians identified five preliminary diagnoses based on vital signs and medical history. The same anonymized data were entered into AI models, which were asked to generate five preliminary diagnoses. After physical examination, laboratory, and imaging results became available, both physicians and AI models refined these into three differential diagnoses. Finally, the physician's final diagnosis, explaining the patient's clinical condition, was accepted as the gold standard. Each AI model was then asked to determine the final diagnosis using the complete dataset. The preliminary, differential, and final diagnoses generated by each model were compared with the gold standard. Statistical analyses were performed using R. Results: The study included 180 patients (56% male, 44% female). Regarding the accuracy of final diagnoses, GPT-4o achieved an accuracy rate of 67.2%, GPT-5 achieved 65.6%, Claude achieved 63.3%, and Gemini achieved 59.4%. No statistically significant difference was found among the groups (p=0.16).When evaluating the agreement between the AI models' definitive diagnoses and the gold-standard diagnosis, Cohen's κ coefficients were calculated as 0.656 for GPT-4o, 0.638 for GPT-5, 0.616 for Claude, and 0.575 for Gemini. Conclusion The finding that AI models achieved diagnostic accuracy rates of 60%-70% and showed overall substantial agreement with the gold standard decision supports their potential to assist clinical decision-making in the ED setting. Keywords: Artificial intelligence, ChatGPT, Claude, Gemini, Large language model

Author

Dr. İbrahim Günaydın

How to Cite

İbrahim Günaydın (Medical Specialty Thesis). Performance of artificial intelligence in critical patient diagnosis: Comparison of ChatGPT, Gemini, Claude and emergency physicians, 2025, Karadeniz Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Karadeniz Technical University