Master'sOpen Access

Multi-class ship image classification using a hybrid vit-resnet architecture with attention mechanisms and semantic gated fusion

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2025
0 views
0 downloads

Abstract (EN)

In this thesis, a hybrid model based on Vision Transformer (ViT) and ResNetRS50 is developed for multi-class classification of ship images. While ViT extracts high-level semantic information, ResNetRS50 captures low- and mid-level spatial features; these two structures are integrated through attention mechanisms and a Gated Fusion layer. During training, advanced techniques such as MixUp and CutMix data augmentation, Focal Loss combined with knowledge distillation loss, the OneCycleLR scheduler, automatic mixed precision (AMP), and exponential moving average (EMA) of model weights are employed. Experiments conducted on a dataset consisting of eight ship classes demonstrate that the proposed architecture outperforms single-stream CNN and ViT models in terms of both accuracy and F1-score. The results indicate that hybrid architectures and attention-based fusion strategies provide an effective solution to the ship classification problem.

Author

Berkay Ergün

How to Cite

Berkay Ergün (Master Thesis). Multi-class ship image classification using a hybrid vit-resnet architecture with attention mechanisms and semantic gated fusion, 2025, Çankaya University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çankaya University