Master'sOpen Access

Turkish tweets on distance education using machine learning methods sentiment analysis

2022
0 views
0 downloads
Advisor: Doç. Dr. İsmail Babaoğlu

Abstract (EN)

The development of technology has led to the development of social media platforms and reaching large user masses. People can communicate with other people by using social media platforms, and they can come together on these platforms to share their feelings and thoughts on a common topic about a product or a topic, in the face of social events that occur. These shares constitute a large data source that can be used in many fields for sentiment analysis studies. With sentiment analysis studies, these data can be processed and analyzed, and positive, negative or neutral emotional expressions about the relevant subject can be determined. With the corona virus epidemic, which started in January 2020 and affected the whole world, some measures were taken across the country, and within the scope of these measures, the distance education process started in March 2020. In this study, Turkish tweets on distance education shared on the social media platform Twitter were obtained and the data were pre-processed and also normalized with the zemberek library and made processable. For the data set to be used as input in the classification phase, besides manual labeling, models were created with TextBlob, Vader and Bert, which directly produce English emotion outputs by making language translation with a different approach. For the data set to be used as input in the classification phase, besides the manual labeling process, a different approach was brought with language translation and models were created with TextBlob, Vader and Bert, which directly produce English emotion outputs. These models were used with different digitization methods (BoW, TF-IDF, Word2Vec,) and different machine learning algorithms (LR, SGD, SVM, RF, NB) and sentiment analysis of the shares made on the best performing classification model was performed. In the structure where Turkish texts have manual tags, 0.79 classification success was achieved with the best TF-IDF – LR pair. When the texts labeled with the manual method and marked as neutral were removed from the data set, it was seen that the success rate increased and the BoW – LR pair gave the best result with a ratio of 0.84. In the models created by labeling ready-made models with the language translation process, the desired level of success for Turkish texts was not achieved.

Author

Dr. Ali Can Akdeniz

How to Cite

Ali Can Akdeniz (Master Thesis). Turkish tweets on distance education using machine learning methods sentiment analysis, 2022, Konya Technical University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Konya Technical University