Detection of offensive language in Turkish tweets using transformer based models
2025
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Sinem Akyol
Abstract (EN)
Social media platforms, which have become one of the most prominent channels of digital communication today, provide individuals with spaces where they can freely express their thoughts, yet at the same time, they also create an environment in which various forms of harmful content can rapidly spread. In particular, expressions that include slang, profanity, and insults can lead to negative consequences at both individual and societal levels, contributing to the prevalence of hate speech, cyberbullying, and digital violence. Within this framework, the automatic analysis of language used in social media, the detection of harmful content, and its filtering when necessary have become essential needs. In this context, artificial intelligence-based natural language processing approaches offer an effective solution, especially in dealing with noisy data sources such as social media texts, which are often filled with abbreviations, colloquialisms, and informal expressions. This thesis aims to automatically detect offensive language in Turkish social media content, and in this scope, a comprehensive comparative analysis was conducted using five different transformer architectures. The evaluated models include BERT, DeBERTa, RoBERTa, XLM-R, and ELECTRA. The main research hypothesis of the thesis is that, under fixed experimental conditions (such as identical training parameters, number of epochs, and preprocessing steps), these models can be compared on the same Turkish dataset in terms of accuracy, training time, and loss metrics to determine which model performs best. The dataset used for this purpose consists of 13,000 Turkish tweets collected from Twitter and manually annotated. During the annotation process, tweets were categorized into two classes: samples labeled as "positive" represent non-offensive, whereas those labeled as "negative" represent offensive content. Within the scope of the study, all models were trained under the same conditions for 35 epochs. Their performances were compared using evaluation metrics such as accuracy rate, training duration, and loss value. According to the findings, the BERT model achieved the highest accuracy with a score of 89.61%. It was followed by the XLM-R, ELECTRA, RoBERTa, and DeBERTa models, respectively. These results indicate that transformer architectures, which surpass traditional deep learning models, can produce successful results even in morphologically rich languages such as Turkish. Moreover, it has been demonstrated that transformer-based approaches have significant potential for the automatic detection of hate speech and profanity, which are commonly encountered in social media.
Author
Zeynep Şebnem Üzmez
How to Cite
Zeynep Şebnem Üzmez (Master Thesis). Detection of offensive language in Turkish tweets using transformer based models, 2025, Fırat University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Fırat University
- Using social media as an integrated marketing communication tool(2018)
- Foundation of Dutch East İndia Company and her rising in İndonesia in the 17th century(2013)
- Examination of stress state between Doğanyol (Malatya) and Çelikhan (Adıyaman) on the east Anatolian fault zone(2020)
- Color usage at Turkish Divan of Fuzûlî(2013)
- Yavuzeli (Gaziantep) surrounding volcanic outcropping of rocks petrographic and geochemical features(2014)
- Hizbu?t-Tahrir and the religions and political thoughts of Ercumend Özkan(2008)