Master'sOpen Access

Feature Selection Using Co-occurrence of Terms for Binary Text Classification

2015
1 views
0 downloads

Abstract (EN)

ABSTRACT: In this thesis, term selection for text categorization is addressed. Three widely used schemes are employed for this purpose, namely Chi-square (x2), Gini_index and Discriminative Power Measure (DPM). The performances of these schemes are evaluated on Reuters-21578 separately for document frequencies and term frequencies. In summary, utilizing the term frequencies leads to better macro and micro F1 score when compared to using only document frequencies. As an extension to the conventionally used term selection schemes, we studied the use of co-occurrence statistics of different terms for feature selection. More specifically, the idea is to evaluate the discriminative power of having two different terms in the selected list at the same time. In order to achieve this, an iterative scheme is designed where the next term to be included in the selected list is determined by pairwise evaluation of the already selected terms and the candidate terms. For the pairwise evaluation of different terms, novel metrics based on the existing selection schemes are developed. Experimental results have shown that the proposed iterative scheme has the potential to improve the existing schemes. Keywords: Term Selection, Text Classification, x2, Gini-index, DPM, Bag-of-Words. …………………………………………………………………………………………………………………………

Author

Dr. Marzieh Vahabi Mashak

How to Cite

Marzieh Vahabi Mashak (Master Thesis). Feature Selection Using Co-occurrence of Terms for Binary Text Classification, 2015, Eastern Mediterranean University, Department of Computer Engineering.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eastern Mediterranean University