DoctorateOpen Access

Türkçe sosyal ortam verileri için makine öğrenme yöntemlerinin geliştirilmesi

2021
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Özlem Aktaş

Abstract (EN)

In this thesis, we have presented a hybrid methodology, which combines the lexicon-based and machine learning (ML)-based approaches for sentiment analysis in Turkish. To use on the lexicon-based side, we have generated a sentiment dictionary by extending SentiTürkNet with a synonym dictionary, ASDICT. Besides this, we have tackled the classification problem with three supervised classifiers, Naive Bayes, Support Vector Machines, and J48, on the ML side. Our hybrid methodology combines these two approaches by generating a new lexicon-based value according to our proposed feature generation algorithm and feeds it as one of the features to ML classifiers. We have experimented on three different datasets such as Movie, Hotel, and Twitter. Despite the linguistic challenges caused by the morphological structure of Turkish, the experimental results show that it improves the accuracy by 7% on average. In conclusion, we have achieved these contributions in our study: It is the first hybrid approach for Turkish sentiment analysis. We have also adapted lemmatization in natural language processing for Turkish SA to preserve the positive and negative meanings of tokens. Finally, we have generated eSTN by extending STN, which is the first comprehensive polarity lexicon for Turkish.

Author

Dr. Buket Erşahin

How to Cite

Buket Erşahin (Doctorate thesis). Türkçe sosyal ortam verileri için makine öğrenme yöntemlerinin geliştirilmesi, 2021, Dokuz Eylül University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Dokuz Eylül University