DoctorateOpen Access

Identification of cyberbullying using machine learning techniques

2024
0 views
0 downloads
Advisor: Prof. Dr. Selma Ayşe Özel

Abstract (EN)

The pervasive utilization of social media platforms has introduced a multitude of threats, among which cyberbullying stands as a significant concern. Cyberbullying is defined as the repetitive use of social and electronic media, to perpetrate immoral actions with the intent to inflict harm upon a victim. In both Turkey and globally, cyberbullying is recognized as a pressing issue necessitating urgent attention due to its profound emotional and psychological repercussions, which can include depression, stress, anxiety, and in extreme cases, suicide. For this concern, numerous researchers and scientists have endeavored to develop solutions leveraging machine learning (ML), deep learning (DL), and natural language processing (NLP) techniques. These efforts aim to create models capable of identifying cyberbullying within textual contexts, with the ultimate goal of either detecting and removing such content or preventing its dissemination. However, the large number of studies in this context have been conducted in English, with a dearth of research in Turkish language. Moreover, existing studies often segregate their analyses between English and Turkish, overlooking potential cross-linguistic nuances. Motivated by these gaps in the literature, this thesis endeavors to address the detection of cyberbullying in both English and Turkish languages. Framed as a binary text classification problem, the research scrutinizes the efficacy of standard techniques across languages, aiming to discern whether a unified approach can effectively detect cyberbullying in diverse linguistic contexts. To achieve this objective, comprehensive experiments are conducted, including the comparative evaluation of ML, DL, and Large Language model (LLM) Furthermore, an array of feature extraction techniques, including traditional feature weightings and word embeddings, are rigorously assessed. Additionally, the efficacy of an optimization technique LoRA, applied to LLM, is thoroughly evaluated. Furthermore, recognizing the problem created by labeled data shortage, a novel semi-supervised learning technique is proposed. Specifically, a self-training with LLM is implemented, demonstrating its potential to enhance classification performance by leveraging a blend of both labeled and unlabeled data, thereby mitigating the resource-intensive nature of manual labeling processes. This research contributes to advancing cyberbullying detection methods and encourages more inclusive approaches across languages to combat cyberbullying in the digital sphere.

Author

Dr. Alı Najıb

How to Cite

Alı Najıb (Doctorate thesis). Identification of cyberbullying using machine learning techniques, 2024, Çukurova University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çukurova University