Master'sOpen Access

Text classification based on organizational data using machine learning

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2019
0 views
0 downloads

Abstract (EN)

The increase in text data coming with the increase in the use of online platforms and the ease of access to this data have led to the increase in the number of studies on text classification. Text classification has had a great impact on such fields as spam mail detection, sentiment analysis and news categorization. Our concern in this study is Turkish text classification. While there are a lot of English papers related to text classification, the number of the studies on Turkish data is quite limited. In this study, the letters of request that came to an organization were used as the experimental dataset. These letters of request are labeled with classes. These classes are predefined in the internal processes of the organization from which the data is received. Because the letters of request used in the study came directly from the users, they contain a lot of misspellings. To correct these mistakes, normalization was applied on the text data. Then, the words in the corpus were transformed into their simple forms by morphological analysis. In addition, the list of stop words was prepared by looking at the most repetitive word groups, and they were removed. Lastly, in the preprocessing step, the repetitive classes in the corpus were simplified via the K-Means algorithm and the number of the classes was reduced. As a result, a more consistent and balanced dataset appeared. The letters of request were trained by Naïve Bayes, SVM, Random Forest, Logistic Regression and LSTM before and after preprocessing. Then, the performance of the algorithms before and after preprocessing was compared. It was concluded that the most efficient algorithm with regard to accuracy is LSTM. Moreover, the model of SaaS was developed for the organizations to benefit from the machine learning model and for increasing the data.

Author

Ahmed Enis Erkaya

How to Cite

Ahmed Enis Erkaya (Master Thesis). Text classification based on organizational data using machine learning, 2019, Ankara Yıldırım Beyazıt University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Ankara Yıldırım Beyazıt University