Feature selection for short text classification
2020
0 views
0 downloads
Advisor: Doç. Dr. Alper Kürşat Uysal
Abstract (EN)
High dimensionality problem is an important concern for short text classification due to its effect on computational cost and accuracy of classifiers. Also, short text data, besides being high dimensional, has an incomplete, inconsistent and sparse structure. Selection of important features that provides a better representation is a solution for high dimensionality problem. However, it is a fact that in feature selection process, short texts need feature selection approaches that will be least affected by the sparse problem. In this study, two feature selection approaches that will work effectively in the short text field are presented for this purpose. Firstly, we developed a novel filter feature selection method called Proportional Rough Feature Selector (PRFS) which uses the rough set for a regional distinction according to the value set of term to identify documents that to be exact belong to a class and have a possibility for belonging to a class. Documents which are possible to belong to a class are penalized by multiplying with a coefficient named α. Additionally, the effect of sparsity in the term vector space is calculated with the help of rough set. The PRFS is compared with state-of-the-art filter feature selection methods such as Gini index, information gain, distinguishing feature selector, recently proposed max-min ratio and normalized difference measure methods. The comparison is carried out using various feature sizes on four different short text datasets with Macro-F1 success measure. Experimental results demonstrated that the PRFS offers either better or competitive performance with respect to other feature selection methods in term of Macro-F1. This study may be a pioneering study in this research field as it proposes a novel feature selection method for short text classification using rough set theory. Secondly, this study presents a new filter feature selection method called XY method which represents the features on XY line and calculates the distance of a feature to the XY line. Also, a value like λ is calculated. According to this value, the terms are divided into different regions such as negative, positive, and third. The XY method aims to select as few terms as possible in the negative region. The XY method is compared with well-known filter feature selection methods such as chi-square, information gain, deviation from Poisson distribution, recently proposed max-min ratio, and distinguishing feature selector methods. The comparison is carried out using various feature sizes in order to make a fair evaluation on four different short text datasets with Macro-F1 success measure. Experimental results demonstrate that the XY method offers either better or competitive performance with respect to other feature selection methods in term of Macro-F1.
Author
Dr. Rasım Çekik
Institution
How to Cite
Rasım Çekik (Doctorate thesis). Feature selection for short text classification, 2020, Eskişehir Teknik Üniversitesi.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Eskişehir Teknik Üniversitesi
- Effect of crystallographic orientation on ionic conductivity of Li(1+x)AlxTi(2-x)(PO4)3 solid electrolytes(2018)
- The investigation of mechanical and dynamic properties of two dimensional mxene crystals by first principles(2018)
- Aircraft sensor fault detection and system reconstruction based on artificial neural networks(2021)
- Fuzzy graphs(2022)
- Production of functionally graded SiC-TiB2-Al composites by spark plasma sintering technique and their characterization(2018)
- Analysis of child mortality with the help of geographic information systems(GIS)(2018)
