Master'sOpen Access

Analysis of non-Latin content on the English information retrieval datasets

2019
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Ahmet Arslan

Abstract (EN)

For centuries people have been aware of the importance of archiving and finding information. With the advent of computers, it is possible to store large amounts of information and finding useful information from such collections became a necessity. The field of Information Retrieval emerged from this requirement in the 1950s. Information retrieval is the process of finding resources that are relevant to an information the users need from large collections. The success of information retrieval systems is directly proportional to the fact that the documents found are related to the information the user is looking for. The Text Retrieval Conference is organized annually to measure the success of information retrieval systems and to compare their performances. Standard data sets are created and published by this organization. In this study ClueWeb09, ClueWeb12 and Gov2 data sets, which consist of English web pages collected from the Internet, are used. Although the majority of the words in these web pages are written in the Latin alphabet, datasets also include words written in non-Latin alphabets (Japanese, Cyrillic, Greek, Arabic, etc). Moreover, the query sets associated with these datasets consist of words written entirely in Latin alphabet. In this context, the objective of this thesis is to examine the distribution of words written in non-Latin alphabets on English data sets and to investigate the effect of including or excluding non-Latin words in index on information retrieval effectiveness.

Author

Dr. Ahmet Alkılınç

How to Cite

Ahmet Alkılınç (Master Thesis). Analysis of non-Latin content on the English information retrieval datasets, 2019, Eskişehir Teknik Üniversitesi.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Eskişehir Teknik Üniversitesi