DoctorateOpen Access

New approaches to enhancing the performance of text classification

2013
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Serkan Günal

Abstract (EN)

The aim of text classification, also known as text categorization, is to classify texts of interest into appropriate classes. Due to the rapid advance of Internet technologies, the amount of electronic documents has drastically increased worldwide. Consequently, text classification has gained importance in organization of these documents. Important issues in text classification are the high dimensionality of feature space and misclassification concerns regarding the feature space. In this dissertation, various solutions are proposed to overcome both of these concerns of the text classification problems. Specifically, a novel filter-based feature selection method, namely distinguishing feature selector, is introduced. Besides, genetic algorithm oriented latent semantic features, which are originated from feature selection and transformation operations, are proposed. Moreover, the impact of several feature extraction and selection approaches on SMS spam filtering problem, a special case of text classification, is extensively investigated for two different languages. Finally, the impact of preprocessing methods on text classification is examined for different domains and different languages as well. Extensive experiments conducted on benchmark datasets revealed that all the proposed solutions offer better dimensionality reduction and/or classification performance depending on their contributions.Keywords: Text Classification, Feature Extraction, Feature Selection, Feature Transformation.

Author

Alper Kürşat Uysal

How to Cite

Alper Kürşat Uysal (Doctorate thesis). New approaches to enhancing the performance of text classification, 2013, Anadolu University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Anadolu University