DoctorateOpen Access

Development of a new model based on local binary and ternary patterns to recognize emotions from speech sounds

2021
0 views
0 downloads
Advisor: Prof. Dr. Asaf Varol

Abstract (EN)

Emotion recognition from speech sounds is an essential discipline which serves to keep the interaction and cooperation between human and computer. It is difficult to analyze the speech signal containing various frequencies and characteristics; thus, speech data-based emotion recognition is a complex problem for machine learning. Even though different methods have been developed for sound/voice/speech feature extraction and classification, success rates vary depending on the languages, emotions, and databases. In this thesis, a new process is proposed that can be applied to databases with different sizes, has low calculation complexity, is low cost, and increases the classification performance. Specifically, a new strategy that makes a contribution to feature extraction technique via reaching the global features from local features is developed. The proposed model consists of three main stages. These stages are feature extraction, feature selection, and classification. Low pass filter coefficients are obtained by applying nine-level One Dimensional Discrete Wavelet Transform (1D-DWT) to raw audio data. Afterward, feature extraction and feature combining are achieved by applying One Dimensional Local Binary Pattern (1D-LBP) and One Dimensional Local Ternary Pattern (1D-LTP) to each of the low pass filters. Local and textural features are obtained by using 1D-LBP and 1D-LTP. Noises in the speech signals are eliminated, the speech signal size is reduced, and new features are extracted through a sequential structure creating nine-level 1D-DWT. In the feature extraction phase, 1D-LBP, 1D-LTP, and 1D-DWT are used together and a new multi-level manual feature extraction process is presented. The most effective features that will be used as inputs to the classifier are selected with distance-based Neighborhood Component Analysis (NCA) while other features are eliminated. Support Vector Machines (SVM), a powerful classifier, is used during the classification phase. The proposed model is tested, without depending on the textual and speaker, in different databases of RAVDESS, EMODB, SAVEE, and EMOVO. Within this framework, in the field of emotion recognition from speech sounds, a new model that increased the classification rating is provided to the literature.

Author

Yeşim Ülgen Sönmez

How to Cite

Yeşim Ülgen Sönmez (Doctorate thesis). Development of a new model based on local binary and ternary patterns to recognize emotions from speech sounds, 2021, Fırat University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Fırat University