Master'sOpen Access

Cyberbullying detection using text classification for turkish language

2019
0 views
0 downloads
Advisor: Prof. Dr. Selma Ayşe Özel ; Dr. Öğr. Üyesi Esra Saraç Eşsiz

Abstract (EN)

Cyberbullying is an electronic form of peer harassment. It includes relational attack behaviors such as harassing people, mocking people, threatening, spreading gossip, and insulting people on the internet by using information and communication technologies. In Turkey and many European countries, the cyberbullying is considered as a serious problem after the cyberbullying related suicides occurred. In recent years, researches are being carried out and solutions are tried to be found by experts, especially with educational scientists and psychologists, about cyberbullying. The aim of this study is to create the largest Turkish dataset so far for the detection of cyberbullying texts and to show the effects of preprocessing, feature selection and classifiers for the detection of cyberbullying from texts. In this study, a number of preprocessing steps are applied, and two well-known filter-based methods that are information gain and chi square are used for feature selection. Among the classifiers tested, Naive Bayes Multinomial is determined to be the most successful method for detecting cyberbullying from texts written in Turkish language. In addition, a filter-based classifier is proposed, and its performance is tested on the collected dataset. The proposed method has promising accuracy and can be used for labeling any Turkish text document without re-training the classifier.

Author

Dr. Erhan Öztürk

How to Cite

Erhan Öztürk (Master Thesis). Cyberbullying detection using text classification for turkish language, 2019, Çukurova University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çukurova University