Development of hybrid systems for convolutional neural networks and vision transformer architectures for medical image datasets
2026
0 views
0 downloads
Advisor: Prof. Dr. Cemil Çolak
Abstract (EN)
Aim: The aim of this study is to comprehensively compare the performance of CNN, ViT, and their hybrid combinations in the classification of thyroid nodule ultrasound images, and to evaluate the effectiveness of a hybrid model (HuviT) that combines the strengths of both approaches. Materials and Methods: A publicly available dataset consisting of 6,905 thyroid ultrasound images was used in the study. The images were resized, normalized, and augmented using various data augmentation techniques. Eleven different deep learning architectures were trained for performance evaluation, including ResNet-50, EfficientNet-B4, ConvNeXt, Swin Transformer, DeiT, PVTv2, XCiT, CoaT, MobileViT, ConViT, and an original hybrid model, HuviT. Performance comparison was conducted using accuracy, precision, recall, F1-score, specificity, and ROC-AUC metrics. To enhance the transparency of model decisions, explainable AI (XAI) methods such as Grad-CAM, Score-CAM, and Attention Rollout were integrated. Results: Experimental results showed that the hybrid architecture ConViT model achieved the highest performance across all metrics (83.30% accuracy, 83.30% precision, 83.26% recall, 83.30% F1-score, 89.94% AUC). ConvNeXt (82.63% accuracy, 90.02% AUC) and DeiT (81.18% accuracy, 89.10% AUC) models also demonstrated high and competitive performance. The original hybrid model HuviT achieved 76.35% accuracy. XAI analyses confirmed that the decision mechanisms of the models were consistent with radiological examination. Conclusion: This study demonstrated that hybrid architectures combining CNN's local feature extraction capability with ViT's global context modeling capacity can achieve superior performance in thyroid ultrasound image analysis compared to pure architectures. The high and balanced success of the ConViT model, in particular, highlights the potential of such hybrid approaches as reliable tools in clinical decision support systems. Keywords: Deep learning, convolutional neural networks, vision transformer, hybrid models, image processing, thyroid nodule classification, explainable artificial intelligence.
Author
Hasan Ucuzal
How to Cite
Hasan Ucuzal (Doctorate thesis). Development of hybrid systems for convolutional neural networks and vision transformer architectures for medical image datasets, 2026, İnönü University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from İnönü University
- Regional threats and opportunities to turkey's national economic security(2022)
- Researching the effect of avenanthramide C on breast cancer(2022)
- The aim of the present study is to examine the etiological origins of cryptogenic cirrhosis in patients who were followed up with the disease(2020)
- Investigation of parents' digital parenting awerness(2020)
- Nutritional monitoring of nutrition in children with cancer(2018)
- Water purification in religions conception of baptism in Christianity(2019)