Text classification based on organizational data using machine learning
Bu tez size mi ait?
Bu kayıt toplu arşivden geldi. Sizinse profilinize bağlayın.
Özet (EN)
The increase in text data coming with the increase in the use of online platforms and the ease of access to this data have led to the increase in the number of studies on text classification. Text classification has had a great impact on such fields as spam mail detection, sentiment analysis and news categorization. Our concern in this study is Turkish text classification. While there are a lot of English papers related to text classification, the number of the studies on Turkish data is quite limited. In this study, the letters of request that came to an organization were used as the experimental dataset. These letters of request are labeled with classes. These classes are predefined in the internal processes of the organization from which the data is received. Because the letters of request used in the study came directly from the users, they contain a lot of misspellings. To correct these mistakes, normalization was applied on the text data. Then, the words in the corpus were transformed into their simple forms by morphological analysis. In addition, the list of stop words was prepared by looking at the most repetitive word groups, and they were removed. Lastly, in the preprocessing step, the repetitive classes in the corpus were simplified via the K-Means algorithm and the number of the classes was reduced. As a result, a more consistent and balanced dataset appeared. The letters of request were trained by Naïve Bayes, SVM, Random Forest, Logistic Regression and LSTM before and after preprocessing. Then, the performance of the algorithms before and after preprocessing was compared. It was concluded that the most efficient algorithm with regard to accuracy is LSTM. Moreover, the model of SaaS was developed for the organizations to benefit from the machine learning model and for increasing the data.
Yazar
Ahmed Enis Erkaya
Bu Yayına Nasıl Atıf Yapılır
Ahmed Enis Erkaya (Master Thesis). Text classification based on organizational data using machine learning, 2019, Ankara Yıldırım Beyazıt University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Ankara Yıldırım Beyazıt University tezlerinden daha fazlası
- Investigation of family functionality detected by adolescents with peer bullying(2019)
- Urban life in Mosul according to the sâlnâmes (1308-1330/1891-1912)(2025)
- Obstacles of e-government development in Yemen(2022)
- Characteristics of patients with epilepsy admitted to the pediatric emergency service(2022)
- Trend networks of Twitter: Examining trends of Twitter Turkey through the concept of network society(2022)
- The impact of the Arab Spring on conflicts in the MENA region: Findings from count data analysis(2022)