DoctorateOpen Access

Various approaches to emotion recognition from speech signals

2020
0 views
0 downloads
Advisor: Doç. Dr. Humar Kahramanlı

Abstract (EN)

Speech emotion recognition from data gains significance in recent years. Emotion recognition has been made from facial expressions and biomedical signals as well as from speech data. Emotion recognition from speech has been used when there is no face-to-face communication. Manually feature extraction and feature selection were the most important steps in traditional speech emotion recognition. Spectral, prosodic and format features are the most frequently extracted features in this area. There were many methods which have been proposed for feature selection. Despite this, the problem of emotion recognition from the speech has not been solved completely and varios techniques are needed to increase the recognition rate. Therefore, in this thesis, emotion recognition from sound was carried out. Berlin Emotion Database (EmoDB), which is the most known and open access database, was used in the study. EmoDB is speech database consist of seven emotion. In this thesis the applications were performed gender and person independent. In this thesis, three different applications were carried out. Two different preliminary studies were carried out in order to explain the data sets and pre-processes to be used in the applications. In the preliminary studies, firstly, the data sets used in Application 1 and Application 2 were created. Data sets are made up of from different numbers and different features. The features of Mel Frequency Kepstrum Coefficients (MFCC), Linear Prediction Coefficients (LPC), Discrete Wavelet Transform (DWT) and Autoregressive Parameters (AR) were extracted with different dimensions. In addition, the format features which were format frequency and pitch have been extracted. ANN, DVM, kNN and NB algorithms are used for classification purposes. In the preliminary studies, secondly the data sets used in Application 3 and pre-processing was explained. First, the Agent Based Automatic Feature Selection (ABAfs) approach was proposed for feature selection. The study applied on the effective data sets, determined the features selected in the in the preliminary studies, were classified. The data sets created in the preliminary study were given to the classifier. The third and last application of the study was emotion recognition with deep learning algorithms. Before performing emotion recognition in the first two applications, feature extraction and feature selection has been performed. In this application, spectrogram images are obtained from raw data without any feature selection. Afterwards it was classified with AlexNET algorithm. In addition, MFCC attributes are classified with DNN in order to compare the classification success of manually extracted features. When all the studies conducted within the scope of the thesis were evaluated, the highest classification accuracy (SD) in seven emotion groups was achieved with a data set consisting of 16 MFCC coefficients with a BCO feature selection method with a rate of 92.98%, and it was added to the literature. In addition, when the literature is examined, the fact that feature selection methods, which have not been applied to emotion recognition until now, have been carried out in this study reveals the originality of the thesis study. So this thesis study will have an important place in the literature in terms of the results obtained, the studies conducted on the determination and classification of which features are more effective in the Emotion Recognition problem.

Author

Dr. Semiye Demircan

How to Cite

Semiye Demircan (Doctorate thesis). Various approaches to emotion recognition from speech signals, 2020, Konya Technical University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Konya Technical University