Speech emotion detection model development using artificial intelligence methods
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
2025
0 views
0 downloads
Advisor: Doç. Dr. Erhan Akbal
Abstract (EN)
Speech emotion recognition is an increasingly significant topic in artificial intelligence, aiming to enable natural and intuitive communication in human-computer interactions. This thesis aims to propose a Turkish speech-to-emotion recognition dataset to the literature and develop a more effective model by combining deep learning-based and traditional methods to identify emotional states from speech data. Within the scope of the thesis, a Turkish speech emotion recognition dataset consisting of 2547 different data with 90 different participants was prepared. The success of the self-supervised learning-based voice models on the prepared dataset was tested and it was shown that the most successful model can achieve more efficient results with the fine-tuning method. In addition, combining multiple datasets will improve the performance of emotion recognition and the selection of the best feature set will significantly improve the model accuracy. In the experimental studies, multiple benchmark datasets comprising individuals with diverse demographic and linguistic characteristics were utilized. These datasets were merge to create a combined dataset consisting of 11,511 speech samples, aiming to improve the generalizability of the proposed model. This study introduces an innovative model designed to enhance emotion performance by combining audio features derived from deep learning-based and traditional methods. The proposed model extract audio features using the Wav2vec2.0 deep learning model and the openSmile audio processing library. Subsequently, the optimal feature set is determined using an iterative feature selection and majority voting approach, which combines the strengths of different feature sets. The experimental results demonstrate that proposed method achieves results comparable to those reported in the existing literature. The proposed approach achieved the highest accuracy of those 92.55% on the multi-dataset corpus. Furthermore, the feature selection and majority voting-based methods proposed in this thesis led to a 3% improvement in classification accuracy. These finding indicate that the combined dataset and selected features significantly enhance model accuracy.
Author
Fatma Güneş Eriş
How to Cite
Fatma Güneş Eriş (Doctorate thesis). Speech emotion detection model development using artificial intelligence methods, 2025, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
- Hizbu?t-Tahrir and the religions and political thoughts of Ercumend Özkan(2008)