Master'sOpen Access

Çizge tabanlı sözcüksel bağdaşıklık ve terim yakınlık heabı ile belge sıralama

2008
0 views
0 downloads
Advisor: Yrd. Doç. Dr. H. Murat Karamüftüoğlu

Abstract (EN)

During the course of reading, the meaning of each word is processed in the context of the meaning of the preceding words in text. Traditional IR systems usually adopt index terms to index and retrieve documents. Unfortunately, a lot of the semantics in a document or query is lost when the text is replaced with just a set of words (bag-of-words). This makes it mandatory to adapt linguistic theories and incorporate language processing techniques into IR tasks. The occurrences of index terms in a document are motivated. Frequently, in a document, the appearance of one word attracts the appearance of another. This can occur in forms of short-distance relationships (proximity) like common noun phrases as well as long-distance relationships (transitivity) defined as lexical cohesion in text. Much of the work done on determining context is based on estimating either long-distance or short-distance word relationships in a document. This work proposes a graph representation for documents and a new matchingfunction based on this representation. By the use of graphs, it is possible to capture both short- and long-distance relationships in a single entity to calculate an overall context score. Experiments made on three TREC document collections showed significant performance improvements over the benchmark, Okapi BM25, retrieval model. Additionally, linguistic implications about the nature and trend of cohesion between query terms were achieved.

Author

Dr. Hayrettin Gürkök

How to Cite

Hayrettin Gürkök (Master Thesis). Çizge tabanlı sözcüksel bağdaşıklık ve terim yakınlık heabı ile belge sıralama, 2008, Bilkent University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Bilkent University