Hate Speech Detection in Social Media
2022
0 views
0 downloads
Advisor: Nazife Dimililer
Abstract (EN)
Hate speech is a phenomenal issue for social media platforms. Recently a rapid increase in hate speech happened all over social media platforms. The aim of this thesis is to improve the performance of the current state-of-the-art for binary text classification in terms of hate speech on social media platforms. The popularity of social media has grown dramatically in recent years. Because of the ease of use and anonymity of the user identity, this increase coincided with the growth of hate speech on social media platforms. Due to the increasing propagation of hate speech, these platforms must implement an automatic hate speech identification system. Hate speech recognition is a difficult task in text mining, due to the use of colloquial language, intentional or incorrect spelling variations. The limitation of the message size on social media platforms also complicates the task since the context of the message is not readily available. Various approaches have been applied to text classification using supervised machine learning models, unsupervised machine learning models, and ensemble approaches. Nevertheless, these approaches did not acquire sufficient confidence to be implemented on social media platforms to address the classification of hate speech. Through this thesis, we proposed two models for detecting hate speech on social media platforms. In the first proposed approach, we developed a model using the novel stacking approach, when two levels of classifiers are used for improving hate speech performance. The second approach based on genetic programming (GP), which is an optimization technique. In the GP approach, a novel mutation technique that combines the standard one-point mutation with a novel feature mutation is employed. Both proposed methods were tested on four publicly available datasets of varying sizes. The experimental results show an improvement in the performance over the other used approaches in this thesis. The results show that the GP approach improves the performance on all datasets, compared to the state-of-the-art in terms of F1-score. On the other hand, in comparison with the state-of-the-art, the stacking approach improves the performance on three over four of the used datasets. Keywords: hate speech, text classification, classifier, classifier ensembles, stacking ensemble, text mining, genetic programming, pattern classification.
Author
Dr. Mona Khalifa A. Aljero
How to Cite
Mona Khalifa A. Aljero (Doctorate thesis). Hate Speech Detection in Social Media, 2022, Eastern Mediterranean University, Department of Mathematics.
Keywords
EN
Applied Mathematics and Computer ScienceComputational intelligenceComputer SecurityComputer securityCyberbullyingHarasssmentHate Speech DetectionHate speechInternetLanguage DetectionMathematicsNatural language processing (Computer science)Social MediaSocial aspectsSpeech Detectionclassifierclassifier ensemblesgenetic programmingpattern classificationstacking ensembletext classificationtext mining
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Eastern Mediterranean University
- An Investigation on Time and Cost Overrun in Construction Projects(2012)
- Radial Power-Law Position-dependent Mass, Cylindrical Coordinates, Spectral Signatures(2015)
- Predicting performance level of reinforced concrete structures subject to corrosion as a function of time(2012)
- Discussion of Conservation Approaches for the Selected Heritage Buildings in the Walled City of Famagusta(2019)
- High School Students' Learning Styles in North Cyprus(2011)
- Afyonkarahisar İl Merkezinde Yaşayan 18 Yaş ve Üzeri Kadınların Diyet Posasıyla İlgili Bilgi Düzeylerinin ve Posa Alım Miktarlarının Belirlenmesi(2018)
