DoctorateOpen Access

Development of hybrid systems for convolutional neural networks and vision transformer architectures for medical image datasets

2026
0 views
0 downloads
Advisor: Prof. Dr. Cemil Çolak

Abstract (EN)

Aim: The aim of this study is to comprehensively compare the performance of CNN, ViT, and their hybrid combinations in the classification of thyroid nodule ultrasound images, and to evaluate the effectiveness of a hybrid model (HuviT) that combines the strengths of both approaches. Materials and Methods: A publicly available dataset consisting of 6,905 thyroid ultrasound images was used in the study. The images were resized, normalized, and augmented using various data augmentation techniques. Eleven different deep learning architectures were trained for performance evaluation, including ResNet-50, EfficientNet-B4, ConvNeXt, Swin Transformer, DeiT, PVTv2, XCiT, CoaT, MobileViT, ConViT, and an original hybrid model, HuviT. Performance comparison was conducted using accuracy, precision, recall, F1-score, specificity, and ROC-AUC metrics. To enhance the transparency of model decisions, explainable AI (XAI) methods such as Grad-CAM, Score-CAM, and Attention Rollout were integrated. Results: Experimental results showed that the hybrid architecture ConViT model achieved the highest performance across all metrics (83.30% accuracy, 83.30% precision, 83.26% recall, 83.30% F1-score, 89.94% AUC). ConvNeXt (82.63% accuracy, 90.02% AUC) and DeiT (81.18% accuracy, 89.10% AUC) models also demonstrated high and competitive performance. The original hybrid model HuviT achieved 76.35% accuracy. XAI analyses confirmed that the decision mechanisms of the models were consistent with radiological examination. Conclusion: This study demonstrated that hybrid architectures combining CNN's local feature extraction capability with ViT's global context modeling capacity can achieve superior performance in thyroid ultrasound image analysis compared to pure architectures. The high and balanced success of the ConViT model, in particular, highlights the potential of such hybrid approaches as reliable tools in clinical decision support systems. Keywords: Deep learning, convolutional neural networks, vision transformer, hybrid models, image processing, thyroid nodule classification, explainable artificial intelligence.

Author

Hasan Ucuzal

How to Cite

Hasan Ucuzal (Doctorate thesis). Development of hybrid systems for convolutional neural networks and vision transformer architectures for medical image datasets, 2026, İnönü University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from İnönü University