Detecting multilingual offensive language in social media using deep neural networks
2023
0 views
0 downloads
Advisor: Prof. Dr. Selma Ayşe Özel
Abstract (EN)
The spread of offensive language on social media platforms has become an alarming reality in society today. The utilization of such language for the purpose of insulting and attacking people represents one of the most detrimental forms of online behavior. Its negative consequences extend to users across different communication platforms, significantly impacting their psychological and mental well-being. To combat this digital malady, data scientists and NLP researchers have taken the task of finding a solution. They have developed several classifier models employing machine learning and deep learning techniques, aimed at identifying several forms of offensive language within textual contexts. These models are designed to process the text by either removing offensive language or preventing its publication on the internet. This study seeks to address the issue by evaluating the performance of deep learning methods on a collected dataset that is formed by collecting a numerous amount of Arabic texts and labeling them. Additionally, comparison of the performance of different deep neural network classifiers namely, Convolutional Neural Network (CNN), Reccurrent Neural Network (RNN), and Long Short-Term Memory (LSTM), and a Language Model namely RoBERTa, performed on the Arabic dataset, as well as some additional datasets in English and Turkish languages, aiming to show the effects of different preprocessing on texts, feature selection and effectiveness of deep neural networks and Transformers across different linguistic texts. The results of this study suggest that RoBERTa is a strong candidate for various language, it achieved the highest validation accuracy across most datasets, showcasing its effectiveness for various languages and tasks. Additionally, an ensemble classifier combining RoBERTa and CNN is introduced and tested, demonstrating good results in improving classification performance.
Author
Dr. Mahmud Birecikli
How to Cite
Mahmud Birecikli (Master Thesis). Detecting multilingual offensive language in social media using deep neural networks, 2023, Çukurova University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Çukurova University
- Subalgebras of free associative algebras(2018)
- Production and characterization of ZnO/Cu2O based devices growing with spin coating method(2019)
- Effect of rations containing black pepper (Piper nigrum) and curcuma (Curcuma Longa Linn) on the performance, egg yield, egg quality properties and blood parameters of hens(2019)
- The effect of bending temperature, holding time, thickness and bending angle on spring back of dual phased high strength steel sheets(2019)
- The rise of populism in liberal world order(2019)
- An evaluation of the performance of the management and development applications employed in wildlife development areas improvement area(2020)
