Detection of deepfake audio and image manuplation with self-supervised learning approach
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Deepfake technology has become a rapidly spreading threat through audio and video manipulations and poses serious security risks. Detection of such fake content is of critical importance in terms of preserving media credibility and preventing the negative effects of fake content on society. In this thesis, it is aimed to detect deepfake video content using two different deep learning-based approaches. In the first approach, a self-supervised method is applied. In this method, pre-training is performed to estimate the rotation angle of the images with a ResNet18-based RotationNet model using the "deepfake-and-real-images" dataset, and then the feature extractor layers of this model are fixed and fine-tuned for real/fake classification. In the second approach, a multi-modality deep learning approach is proposed. In the study, videos taken from the widely used DeepFake Detection Challenge dataset are used. In the proposed method, an MTCNN-based face detection and alignment method is used to obtain face images from video streams. These face images are then given as input to the CNN and LSTM-based model. Mel spectrograms extracted from audio streams were analyzed with a CNN-based model. An audiovisual fusion model was developed that processes features derived from both modalities and their combination. Predictions from the image, audio, and fusion models were combined with learnable weights to create an ensemble model, and the final decision was made based on this model. As a result of the experimental studies, it was observed that the developed self-supervised approach detected deepfake videos with 88.89% accuracy, 83.15% recall and 88.10% F1 score. It was observed that the multimodal model could successfully detect 92.17% accuracy, 90.36% recall and 92.02% F1 score. These studies aim to contribute to the studies in this field by revealing the effectiveness of the methods used in deepfake detection.
Author
Merve Yıldırım
Institution
How to Cite
Merve Yıldırım (Master Thesis). Detection of deepfake audio and image manuplation with self-supervised learning approach, 2025, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
- Hizbu?t-Tahrir and the religions and political thoughts of Ercumend Özkan(2008)