Comparing audio features for speech emotion recognition using machine learning algorithms
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Voice is an integral part of our lives. The demand for voice technology in both art and human-machine interaction systems has recently been increased. More information can be transferred quickly by voice. Speech is a natural way of communicating and as a result of this, it is primarily preferred for contacting users in technological areas. Our voice conveys both linguistic and paralinguistic messages in the course of speaking. The paralinguistic part, for example, rhythm and pitch, provides emotional cues to the speaker. Emotions consist of cognitive, physiological and behavioural changes and all these phenomena are interrelated. Generally, an emotion is a state that affects the thoughts and is capable of determining behaviour. Emotion also creates physical and psychological changes. Speech Emotion Recognition topic examines the question 'How is it said?' and an algorithm detects the emotional state of the speaker from an audio record. Within the scope of this study, machine learning models are developed with classification methods to resolve the problem of speech emotion recognition. Voice consists of a lot of characteristics. However, the optimal audio feature set related to the emotional state cannot be determined yet. The main aim in this study is obtaining the most distinctive emotional features. For this purpose, in order to compare audio features based on different domains Root Mean Square Energy (RMSE), Zero Crossing Rate (ZCR), Chroma and Mel Frequency Cepstral Coefficients (MFCC) features are examined for emotion recognition. A pre-trained model namely wav2vec Large which has been developed more recently is used to create the inputs also. Support Vector Machine, Multi-Layer Perceptron and Convolutional Neural Network techniques are utilized for developing learning models for comparing traditional features and the pre-trained model representations. In this paper emotions namely, Happy, Calm, Angry, Boredom, Disgust, Fear, Neutral, Sad and Surprise are classified, and furthermore, the models are trained and tested with English and German speech datasets. When the classification results are examined, it is concluded that the most successful predictions are obtained with the pre-trained representations. The weighted accuracy ratio is 91% for both Convolutional Neural Network and Multilayer Perceptrons models while this ratio is 87% for the Support Vector Machine models. Among the emotional states, Fear has the highest recognition ratio with 95% f-score with Convolutional Neural Network technique which uses a pre-trained model.
Author
Fatma Gümüş
Institution
How to Cite
Fatma Gümüş (Master Thesis). Comparing audio features for speech emotion recognition using machine learning algorithms, 2022, MEF University.
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from MEF University
- The impact of smartphone use on academic achievement in the digital age: The mediating role of self-regulation(2025)
- The crime of sexual intercourse with minors and its effects on the victim(2025)
- The legal and structural framework of lma-type model loan agreements secured by export credit agencies(2025)
- Eviction due to two justified warnings in residence and roofed workplace rents(2025)
- Design of complex-geometry parts for multi-axis robotic additive manufacturing technology and its simulation(2025)
- The evaluation of Non-Fungible Tokens (NFTs) within the framework of the law on intellectual and artistic works(2025)