Uzun süreli kısa kullanarak konuşma sentezi bellek ve tekrarlamalı sinir ağları (LTSM-RNN)
2023
0 views
0 downloads
Advisor: Prof. Dr. Galip Cansever
Abstract (EN)
Huge advancements in the development of artificial intelligence techniques have been made in the last decade, which have led to the diffusion and spread of computer generated multimedia content, consisting of images, audio and video, which is so realistic that it makes it difficult to be told apart from original content of the same nature. While there are interesting applications to artificial intelligence generated content, it can also be used in dangerous and deceiving ways, for example as proof in a court of law. Hence it is more and more urgent to find automatic ways to distinguish artificial intelligence speech synthesized content from original content. In this research, we take into account audio content, and in particular speech, which is obviously utterly delicate as it comes to forgery. We are going to deepen the previous research made in the field of bi-spectral analysis in order to create more general automatic methods to recognize real speakers from artificial intelligence synthesized speech using Long-Term-Short Memory Recurrent Neural Network (LTSM-RNN). The dataset of voices that we have used is very wide and heterogeneous, consisting of both real voices and voices synthesized using various different methods. We extracted the bi-coherence from all the speech recordings and performed some classifications (both multi-label classifications, which consist in distinguishing each class of voices from all the others, and binary classifications between real and fake voices) using various machine learning and deep learning techniques, such as support vector machine, logistic regression and convolutional neural network. In particular, once the bi-coherences have been computed from the audio files, we performed the following tests. First of all, we replicated the ix tests made on previous works extracting from the bi-coherences a set of features which consist on mean, variance, skewness and kurtosis of both the modules and the phases of the bicoherences and trying to classify them performing simple multiclass and binary classifications using a LTSM, a series of RNN and some CNN. Then we simulated an open set environment using a series of LTSM, in order to test the model with data not yet seen in the training phase. Moreover, we used a series of hybrid LTSM-RNN to extract a new set of features and tried to classify them performing simple multi-label and binary classifications. Finally, research is concatenated the two set of features above and performed more classifications with them (in this case also with an open set environment) and we are going to show that with this method we obtained the best results with an accuracy of 99.76%. The results clarified better the role of bispectral analysis in distinguishing between real and fake speech recordings or synthesis, and lead to more research in the field of multimedia forensics.
Author
Dr. Arkan Adnan Imran Al-yasarı
Institution

Altınbaş University
Elektrik ve Bilgisayar Mühendisliği Bilim Dalı
How to Cite
Arkan Adnan Imran Al-yasarı (Master Thesis). Uzun süreli kısa kullanarak konuşma sentezi bellek ve tekrarlamalı sinir ağları (LTSM-RNN), 2023, Altınbaş University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Altınbaş University
- Mahmutbey, İstanbul'da sosyal dayanıklılık ve toplumsal uyumun güçlendirilmesi(2025)
- Evaluation of the factors affecting the choice of child oral care products and the attitudes of parents to these products(2023)
- Poliüre kaplamanın alüminyum köpük ve katkılı üretilen numunelerin mekanik özelliklerine etkisi(2021)
- Internationalism and a socialist workers' organization in Ottoman Empire: The socialist workers' federation of thessaloniki (1908 - 1914)(2019)
- Symmetry-based multi-objective AI/ML driven optimization framework for sustainable building performance(2026)
- The effect of music and aromatherapy on dental anxiety and fear in children(2024)