Master'sOpen Access

Sentiment analysis in imbalanced datasets

2021
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Aysun Güran

Abstract (EN)

In this thesis, in the subject of sentiment analysis, which has become popular in recent years and its analysis is very important for understanding customer feedback, the problems caused by unbalanced data sets when working with real life data are discussed. In order to get reliable and consistent results on this issue, the methods that are the most recommended solution methods in the literature, based on sample increase and sample reduction were examined. Balancing performances of data sets of these methods and their positive and negative effects on the results have been demonstrated by controlled experiments. Within the scope of the study, 3 different sentiment analysis data sets consisting of social media data compiled from daily life and both Turkish and English text data were studied. Using word-based N-gram structures on the data sets, ROS and SMOTE algorithms for oversampling and RUS and NM1 algorithms for undersampling were applied. Then, the logistic regression classifier and support vectors machine algorithms were analyzed by observing in their effects. As a result, by comparing the observations of the sampling methods for logistic regression, it was observed that increasing the performance was achieved for all N-gram test values for the RUS and SMOTE methods. A similar performance occurred only for certain data sets for support vector machines. For the RUS and NM1 algorithms used in undersampling methods, it was found that the performance value for both logistic regression and support vector machines for all used N-grams decreased and the results were explained in detail.

Author

Hamdi Atacan Oğul

How to Cite

Hamdi Atacan Oğul (Master Thesis). Sentiment analysis in imbalanced datasets, 2021, Doğuş University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Doğuş University