DoctorateOpen Access

Selective stemming for robust information retrieval

2022
0 views
0 downloads
Advisor: Doç. Dr. Ahmet Arslan ; Prof. Dr. Bekir Taner Dinçer

Abstract (EN)

Stemming is supposed to improve the average performance of an information retrieval system, but in practice, past experimental results show that this is not always the case. In this thesis, a selective approach to stemming that decides whether stemming should be applied or not on a query basis is proposed. Our method aims at minimizing the risk of failure caused by stemming in retrieving semantically related documents. The thesis contributes to the information retrieval literature by proposing an application of selective stemming and introducing a set of new features derived from the term frequencies of the words and their stems produced by different stemming algorithms. The proposed method leverages both some of the existing query performance predictors and the newly derived features, and a supervised classification (machine learning) technique. It is evaluated using three rule-based stemmers and eight query sets of the standard TREC and NTCIR tracks. The document collections used, except WSJ document collection, consist of Web documents ranging in number from 25 million to 733 million. The results of the experiments show that the method is capable of making accurate selections that increase the robustness of the system and minimize the risk of failure (i.e., per-query performance losses) across queries. The results also show that the method attains systematically higher average retrieval performance than the single systems for most query sets.

Author

Gökhan Göksel

How to Cite

Gökhan Göksel (Doctorate thesis). Selective stemming for robust information retrieval, 2022, Eskişehir Technical Üniversity.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eskişehir Technical Üniversity