Master'sOpen Access

Detecting Hate Speech using Machine Learning and Sampling Techniques

2021
0 views
0 downloads
Advisor: Nazife Dimililer

Abstract (EN)

The spread of hate speech on social media platforms is a problem that is constantly becoming more imminent as the access to related technologies gets easier. This study focuses on detecting hate speech on an imbalanced multiclass twitter dataset using Machine Learning (ML) algorithms. The most commonly used ML algorithms namely, Logistic Regression, Support Vector Machines (SVM) and deep learning systems Gated Recurrent Unit (GRU), Convolutional Neural Network (CNN), Long Short-Term Memory (LSTM), Bi-directional Long Short-Term Memory (BiLSTM) and a hybrid model CNNBiLSTM have been used for hate speech detection. In order to overcome the problems that arise from using an imbalanced dataset several techniques are used to balance the dataset, Synthetic Minority Oversampling Technique (SMOTE), SMOTETomek, SMOTEENN, Adaptive Synthetic (ADASYN), class weights and the proposed method. Each classifier was trained with all data balancing techniques and their performances were compared in order to find the best classifier for classifying hate speech in the dataset. The best classifier was CNN using the proposed method and it had an F1-score of 0.96 with a Cohen Kappa score of 0.94 and an overall Recall and Precision score of 0.96. For the best system, the recall and precision scores for the hate class was 1.00 and 0.94 respectively.

Author

Dr. Angela Yeukai Chikova

How to Cite

Angela Yeukai Chikova (Master Thesis). Detecting Hate Speech using Machine Learning and Sampling Techniques, 2021, Eastern Mediterranean University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eastern Mediterranean University