Detecting offensive language from social media using word embedding and language models
2023
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Selma Ayşe Özel
Özet (EN)
This research addresses the pressing challenges posed by the proliferation of abusive content on social media platforms, tackling this issue in both English and Arabic languages. To construct a robust framework for detecting offensive language, we have employed cutting-edge methodologies. These include leveraging prominent language models such as Base BERT, Mini BERT, and GPT-2, as well as utilizing LSTM (Long Term Memory) models and an SVM (Support Vector Machine) classifier. Additionally, we have harnessed word embedding techniques like GloVe and Word2Vec to capture the intricate semantic relationships among words. The primary objective of this research is to fortify the detection mechanisms for offensive language, thereby nurturing a safer online environment, especially for vulnerable user groups like children and adolescents. Despite the relatively limited availability of Arabic resources for identifying offensive language, our research bridges this gap. It makes a substantial contribution to the field by presenting an extensive dataset encompassing various Arabic dialects. Through meticulous evaluation, we have optimized the synergy among the aforementioned methods to achieve precise classification of offensive content. In summary, this research aspires to cultivate a safer digital society and deepen our comprehension of the dynamics of offensive language within both Arabic and English social media spheres. Notably, in the English language, our best accuracy of 93.29% was achieved with the HateBERT model and Base BERT in tandem with an SVM classifier, employing a dropout rate of 0.4. For Arabic language, the highest accuracies were attained with the truncated dataset, achieving an accuracy of 89.35% when utilizing AraBERT Tweet, and with the entire dataset using AraBERT Tweet, reaching an accuracy of 92.86%. These achievements mark significant milestones in our pursuit of effective offensive language detection.
Yazar
Dr. Raghad Birecikli
Bu Yayına Nasıl Atıf Yapılır
Raghad Birecikli (Master Thesis). Detecting offensive language from social media using word embedding and language models, 2023, Çukurova University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Çukurova University tezlerinden daha fazlası
- Credit risk management in banking sector: An application of variables determining credit risk in Turkish banking sector(2011)
- Comparasion of the shear bond strength of two different precoated and uncoated ceramic brackets(2014)
- The control tests of four anode photomultiplier tubes for hf calorimeter of CMS detector(2014)
- The predictive strength of career decision making difficulties on high school students' career maturity accordi̇ng to their levels of focus of control(2017)
- Association of heat shock protein with some physiological parameters in the goats(2018)
- Efficacy of thoracic ultrasound in patients presenting to the emergency department with shortness of breath(2019)
