DoctorateOpen Access

Feature selection for text classification and the effect of globalisation

2021
0 views
0 downloads
Advisor: Doç. Dr. Alper Kürşat Uysal

Abstract (EN)

Nowadays, with the increase of internet services, textual data increases exponentially with each passing day. In order to make these texts more meaningful and useful, the texts should be classified according to their content. For this reason, automatic text classification approaches have gained importance. The main task of text classification approaches is to assign texts to classes according to their content. There are many steps to assign text-containing documents to classes suitable for their content. These are feature extraction, feature selection, feature weighting and classification processes. In order to increase the text classification performance, each of these stages has a special importance. However, feature selection has become more popular in recent years. In this thesis, performances were compared using different globalisation techniques (maximum, sum, weighted sum) on local feature selection methods used for text classification and a novel feature selection method with higher performance than the current feature selection methods in the literature are proposed. For this purpose, we have observed how globalisation techniques change performance on datasets with different characteristics. Also, considering the corpus-based and class-based scores of the feature, a new feature selection method is proposed, called Extensive Feature Selector(EFS).

Author

Dr. Bekir Parlak

How to Cite

Bekir Parlak (Doctorate thesis). Feature selection for text classification and the effect of globalisation, 2021, Eskişehir Teknik Üniversitesi.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eskişehir Teknik Üniversitesi