DoctorateOpen Access

Speech emotion detection model development using artificial intelligence methods

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2025
0 views
0 downloads
Advisor: Doç. Dr. Erhan Akbal

Abstract (EN)

Speech emotion recognition is an increasingly significant topic in artificial intelligence, aiming to enable natural and intuitive communication in human-computer interactions. This thesis aims to propose a Turkish speech-to-emotion recognition dataset to the literature and develop a more effective model by combining deep learning-based and traditional methods to identify emotional states from speech data. Within the scope of the thesis, a Turkish speech emotion recognition dataset consisting of 2547 different data with 90 different participants was prepared. The success of the self-supervised learning-based voice models on the prepared dataset was tested and it was shown that the most successful model can achieve more efficient results with the fine-tuning method. In addition, combining multiple datasets will improve the performance of emotion recognition and the selection of the best feature set will significantly improve the model accuracy. In the experimental studies, multiple benchmark datasets comprising individuals with diverse demographic and linguistic characteristics were utilized. These datasets were merge to create a combined dataset consisting of 11,511 speech samples, aiming to improve the generalizability of the proposed model. This study introduces an innovative model designed to enhance emotion performance by combining audio features derived from deep learning-based and traditional methods. The proposed model extract audio features using the Wav2vec2.0 deep learning model and the openSmile audio processing library. Subsequently, the optimal feature set is determined using an iterative feature selection and majority voting approach, which combines the strengths of different feature sets. The experimental results demonstrate that proposed method achieves results comparable to those reported in the existing literature. The proposed approach achieved the highest accuracy of those 92.55% on the multi-dataset corpus. Furthermore, the feature selection and majority voting-based methods proposed in this thesis led to a 3% improvement in classification accuracy. These finding indicate that the combined dataset and selected features significantly enhance model accuracy.

Author

Fatma Güneş Eriş

How to Cite

Fatma Güneş Eriş (Doctorate thesis). Speech emotion detection model development using artificial intelligence methods, 2025, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University