Detecting offensive language from social media using word embedding and language models
2023
0 views
0 downloads
Advisor: Prof. Dr. Selma Ayşe Özel
Abstract (EN)
This research addresses the pressing challenges posed by the proliferation of abusive content on social media platforms, tackling this issue in both English and Arabic languages. To construct a robust framework for detecting offensive language, we have employed cutting-edge methodologies. These include leveraging prominent language models such as Base BERT, Mini BERT, and GPT-2, as well as utilizing LSTM (Long Term Memory) models and an SVM (Support Vector Machine) classifier. Additionally, we have harnessed word embedding techniques like GloVe and Word2Vec to capture the intricate semantic relationships among words. The primary objective of this research is to fortify the detection mechanisms for offensive language, thereby nurturing a safer online environment, especially for vulnerable user groups like children and adolescents. Despite the relatively limited availability of Arabic resources for identifying offensive language, our research bridges this gap. It makes a substantial contribution to the field by presenting an extensive dataset encompassing various Arabic dialects. Through meticulous evaluation, we have optimized the synergy among the aforementioned methods to achieve precise classification of offensive content. In summary, this research aspires to cultivate a safer digital society and deepen our comprehension of the dynamics of offensive language within both Arabic and English social media spheres. Notably, in the English language, our best accuracy of 93.29% was achieved with the HateBERT model and Base BERT in tandem with an SVM classifier, employing a dropout rate of 0.4. For Arabic language, the highest accuracies were attained with the truncated dataset, achieving an accuracy of 89.35% when utilizing AraBERT Tweet, and with the entire dataset using AraBERT Tweet, reaching an accuracy of 92.86%. These achievements mark significant milestones in our pursuit of effective offensive language detection.
Author
Dr. Raghad Birecikli
How to Cite
Raghad Birecikli (Master Thesis). Detecting offensive language from social media using word embedding and language models, 2023, Çukurova University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Çukurova University
- The method of annotation Ibn Allan's work called Dalil al-Falihin(2022)
- Investigation of ruminaton levels of high school students according to parental attitudes and genders(2024)
- Behavioral economics and public policies(2024)
- Çukurova bölgesinde Bisalbuminemi(1977)
- Socio-economic factors influencing immigrant families' from rural areas to Adana in terms of health behavior and attitudes(1979)
- Diabetik sistopatide sistometri(1983)
