DoctorateOpen Access

Deep learning-based multi-class and multi-label classification of ocular diseases using optical coherence tomography and retinal fundus images

2025
0 views
0 downloads
Advisor: Prof. Dr. Abdulkadir Şengür

Abstract (EN)

The classification of multi-class and multi-label images using Artificial Intelligence (AI) techniques remains a significant and actively researched problem in the literature. In this thesis, the performance of deep learning models based on Convolutional Neural Networks (CNN) and Hybrid Vision Transformer (HViT), developed for the classification of ocular diseases using Optical Coherence Tomography (OCT) and fundus images in the field of medical imaging, is evaluated on multi-class and multi-label datasets. Comprehensive experimental studies were conducted on various datasets containing OCT and fundus images to analyze the generalization capabilities of the proposed models and their comparative performance against existing methods in the literature. Experimental studies were carried out on the Kermany UCSD, OCTDL, and OCTID datasets comprising OCT images, as well as the ODIR-5K dataset containing fundus images. The OCT datasets consist of multiclass disease categories, while the ODIR-5K dataset includes both multi-class and multi-label disease annotations. To address the class imbalance problem often encountered in such datasets, various strategies were developed and implemented, including the generation of synthetic images using Generative Adversarial Networks (GANs). Using the proposed CNN-based approach on the Kermany UCSD dataset, the following classification performance metrics were achieved: 96.5% accuracy, 98.29% F1-score, 99.56% AUC, 98.31% precision, and 98.3% recall. In comparison, the proposed HViT model achieved 99.7% accuracy, 99.7% F1- score, 100% AUC, 99.9% precision, 99.7% recall, and a Kappa score of 0.99 on the same dataset. On the OCTDL dataset, the CNN-based method yielded 95.26% accuracy, 92.16% F1-score, 99.74% AUC, 93.11% precision, and 91.53% recall. The HViT model achieved 96.58% accuracy, 96.06% F1-score, 99.92% AUC, 96.65% precision, 95.72% recall, and a Kappa score of 0.86. On the OCTID dataset, the CNN-based model achieved 93.81% accuracy, 92.49% F1-score, 99.96% AUC, 93.37% precision, and 92.23% recall, while the HViT model achieved 94.41% accuracy, 94.14% F1-score, 99.55% AUC, 94.02% precision, 94.4% recall, and a Kappa score of 0.82. For the ODIR-5K fundus image dataset, the proposed HViT model achieved 81.14% accuracy, 81.03% F1-score, 97.99% AUC, 81.08% precision, 81.16% recall, and a Kappa score of 0.80. The obtained results were compared with those reported in the literature, demonstrating that the proposed models achieve high accuracy, F1-score, Kappa score, precision, recall, and AUC in OCT and fundus image classification tasks. Furthermore, it was concluded that factors such as class imbalance, data diversity, and model generalizability should be taken into account. The proposed CNN and HViT models show strong potential as reliable and effective solutions for clinical decision support systems.

Author

Muhammed Halil Akpınar

How to Cite

Muhammed Halil Akpınar (Doctorate thesis). Deep learning-based multi-class and multi-label classification of ocular diseases using optical coherence tomography and retinal fundus images, 2025, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University