Optimized deep learning approach for image augmentation and classification using generative adversarial network and vision transformer
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
2025
0 views
0 downloads
Advisor: Doç. Dr. Kemal Adem
Abstract (EN)
The availability of large, diverse, and balanced datasets often enables the effective applications of deep learning models in image classification tasks. This limitation becomes particularly evident in domain-specific problems where collecting a sufficient amount of labeled image data is challenging. To address this issue, this study proposes an integrated approach that combines synthetic data generation using a Deep Convolutional Generative Adversarial Network (DCGAN) with classification using a Vision Transformer (ViT) model, further enhanced through meta-heuristic hyperparameter optimization. First, the DCGAN model was optimized to generate high-quality and diverse synthetic images, augmenting the limited FSO turbulence dataset and mitigating issues related to data scarcity and imbalance. The quality of the generated images was quantitatively evaluated using Inception Score (IS) and Fréchet Inception Distance (FID) metrics. The Adagrad optimizer combined with the LeakyReLU activation function yielded highly competitive results, achieving an IS value of 1.2342 and the lowest FID value of 67.7493. Next, the classification process was performed by optimizing both the transformer backbone and the CNN-based classification head. Six nature-inspired meta-heuristic algorithms were employed for this purpose. The ViT-HHO and ViT-ABC models achieved the highest overall performance. ViT-HHO attained the highest accuracy of 93.33%. Similarly, ViT-ABC achieved an accuracy of 93.06%. Experimental results demonstrate that the proposed framework significantly enhances classification performance compared to baseline ViT configurations and state-of-the-art transfer learning models, including EfficientNetB7, ResNet-50, DenseNet121, and InceptionV3. The optimized ViT models, trained on GAN-augmented datasets, achieved high accuracy, precision, recall, and F-score metrics, confirming the effectiveness of the integrated data augmentation and optimization strategy. This study highlights the potential of combining GAN-based data generation and ViT-based classification with meta-heuristic optimization to address the dual challenges of limited data and model performance in specialized computer vision tasks.
Author
Emre Yüksek
Institution
How to Cite
Emre Yüksek (Doctorate thesis). Optimized deep learning approach for image augmentation and classification using generative adversarial network and vision transformer, 2025, Sivas University of Science and Technology.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sivas University of Science and Technology
- Control of surface wettability in polymethylmethacrylate films(2026)
- Effects of different endophytic bacteria on in vitro growth of fenugreek (trigonella foenum-graecum)(2026)
- Investigation of the effects of permanent magnet synchronous machines on the power system parameters within the scope of microgrid(2026)
- Smart contact lenses: A literature review on dual-use scenarios in healthcare and defense industry(2026)
- Investigation of the joining of additively manufactured aluminum to copper by friction stir welding(2026)
- 3D profiler with led source and 2D continuous wavelet transform(2026)