Master'sOpen Access

Classification of scientific manuscripts using text processing methods

2013
0 views
0 downloads
Advisor: Doç. Dr. Hasan Şakir Bilge

Abstract (EN)

Transferring of paper-based texts to digital media has become easier with today?s technological advances. Classification of texts should be made in order to access information more easily. Before classification, text processing techniques must be applied many natural language texts. Text processing is the process of analyzing with variety of techniques in order to classify raw data in documents. In this study, a data set of scientific articles published in Turkish was built and it is aimed to obtain high success by applying different text processing and classification methods. With this aim text classification procedures (preprocessing, indexing, feature selection, classification and performance evaluation) were performed step by step. We used character 2-gram and 3-gram methods to choose the word stem in order to express the texts used in this study. To quantintify the data obtained from abovementioned methods, we applied TF, binary and most commonly used TF-IDF weighting methods of the vector space model. We used information gain and correlation based feature selection methods in order to choose the relevant features and remove the unnecessary ones. We used the most famous classifications methods, namely K-NN, Naive Bayes, Multinominal Naive Bayes and SVM, on the Weka software to benchmark the performance of the proposed method. In advance, data set was compared to an other one (1150 news published in Turkish in Internet). In conclusion, the best results regarding the feature vectors obtained using word stems were obtained from the double weighting method. For the character 2-gram and 3-gram methods, the best results were obtained from TF weighting method. The information gain method returned better results compared to the correlation based feature selection method. It yielded better performance on the fusion at feature level. The best result (99,44%) was obtained from the word stems+information gain+binary+TF+TF-IDF feature vector. By applying the text processing methods explained in this study, we obtained better results compared to the previous study.

Author

Samal Kaliyeva

How to Cite

Samal Kaliyeva (Master Thesis). Classification of scientific manuscripts using text processing methods, 2013, Gazi University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Gazi University