Master'sOpen Access

Effects of feature extraction techniques on classification of turkish texts

2018
0 views
0 downloads
Advisor: Prof. Dr. Selma Ayşe Özel

Abstract (EN)

The purpose of this thesis is to determine the most effective method for extracting features and to develop an effective method of selecting features for the classification of Turkish documents in different types. We analyze the effects of preprocessing methods, weighting schemes, and feature selection on the performance of Turkish document classification. In the study, 5 different term weighting methods that are tf, tp, logtf, normtf, tf*idf are compared and it is found that "tf" and "tf*idf" give the best results. After that effects of stopwords removal are investigated, and it is observed that stopwords removal improves classification performance. Then we compare 5 different stemming algorithms that are Zemberek, Affix Stripping, Fixed Prefix 3, 5, and 7 to find out the effects of stemming algorithms on the classification. The results of classification obtained from applying stemming and using the raw form of terms are compared, and the raw form of terms gives more accurate classification results. The effects of n-gram based feature extraction, and feature selection methods that are our proposed standard deviation based method, well-known information gain, and chi-square algorithms are compared. The experimental results indicate that the n-gram feature extraction and standard deviation-based feature selection algorithms give the best results and these methods improve the classification accuracy positively.

Author

Dr. Özge Akdoğan

How to Cite

Özge Akdoğan (Master Thesis). Effects of feature extraction techniques on classification of turkish texts, 2018, Çukurova University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Çukurova University