DoktoraAçık Erişim

Paralel metin getirme için modeller ve algoritmalar

2006
0 görüntülenme
0 i̇ndirme
Danışman: Prof.dr. Cevdet Aykanat

Özet (EN)

ABSTRACTMODELS AND ALGORITHMS FORPARALLEL TEXT RETRIEVALBerkant Barla CambazoğlugPh.D. in Computer EngineeringSupervisor: Prof. Dr. Cevdet AykanatJanuary, 2006In the last decade, search engines became an integral part of our lives. The cur-rent state-of-the-art in search engine technology relies on parallel text retrieval.Basically, a parallel text retrieval system is composed of three components: acrawler, an indexer, and a query processor. The crawler component aims to lo-cate, fetch, and store the Web pages in a local document repository. The indexercomponent converts the stored, unstructured text into a queryable form, mostoften an inverted index. Finally, the query processing component performs thesearch over the indexed content. In this thesis, we present models and algo-rithms for efficient Web crawling and query processing. First, for parallel Webcrawling, we propose a hybrid model that aims to minimize the communicationoverhead among the processors while balancing the number of page download re-quests and storage loads of processors. Second, we propose models for document-and term-based inverted index partitioning. In the document-based partitioningmodel, the number of disk accesses incurred during query processing is minimizedwhile the posting storage is balanced. In the term-based partitioning model, thetotal amount of communication is minimized while, again, the posting storageis balanced. Finally, we develop and evaluate a large number of algorithms forquery processing in ranking-based text retrieval systems. We test the proposedalgorithms over our experimental parallel text retrieval system, Skynet, currentlyrunning on a 48-node PC cluster. In the thesis, we also discuss the design andimplementation details of another, somewhat untraditional, grid-enabled searchengine, SE4SEE. Among our practical work, we present the Harbinger text clas-sification system, used in SE4SEE for Web page classification, and the K-PaToHhypergraph partitioning toolkit, to be used in the proposed models.Keywords: Search engine, parallel text retrieval, Web crawling, inverted indexpartitioning, query processing, text classification, hypergraph partitioning.iv

Yazar

Dr. Berkant Barla Cambazoğlu

Bu Yayına Nasıl Atıf Yapılır

Berkant Barla Cambazoğlu (Doctorate thesis). Paralel metin getirme için modeller ve algoritmalar, 2006, Bilkent University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Bilkent University tezlerinden daha fazlası