Paralel metin getirme için modeller ve algoritmalar
2006
0 görüntülenme
0 i̇ndirme
Danışman: Prof.dr. Cevdet Aykanat
Özet (EN)
ABSTRACTMODELS AND ALGORITHMS FORPARALLEL TEXT RETRIEVALBerkant Barla CambazoğlugPh.D. in Computer EngineeringSupervisor: Prof. Dr. Cevdet AykanatJanuary, 2006In the last decade, search engines became an integral part of our lives. The cur-rent state-of-the-art in search engine technology relies on parallel text retrieval.Basically, a parallel text retrieval system is composed of three components: acrawler, an indexer, and a query processor. The crawler component aims to lo-cate, fetch, and store the Web pages in a local document repository. The indexercomponent converts the stored, unstructured text into a queryable form, mostoften an inverted index. Finally, the query processing component performs thesearch over the indexed content. In this thesis, we present models and algo-rithms for eï¬cient Web crawling and query processing. First, for parallel Webcrawling, we propose a hybrid model that aims to minimize the communicationoverhead among the processors while balancing the number of page download re-quests and storage loads of processors. Second, we propose models for document-and term-based inverted index partitioning. In the document-based partitioningmodel, the number of disk accesses incurred during query processing is minimizedwhile the posting storage is balanced. In the term-based partitioning model, thetotal amount of communication is minimized while, again, the posting storageis balanced. Finally, we develop and evaluate a large number of algorithms forquery processing in ranking-based text retrieval systems. We test the proposedalgorithms over our experimental parallel text retrieval system, Skynet, currentlyrunning on a 48-node PC cluster. In the thesis, we also discuss the design andimplementation details of another, somewhat untraditional, grid-enabled searchengine, SE4SEE. Among our practical work, we present the Harbinger text clas-siï¬cation system, used in SE4SEE for Web page classiï¬cation, and the K-PaToHhypergraph partitioning toolkit, to be used in the proposed models.Keywords: Search engine, parallel text retrieval, Web crawling, inverted indexpartitioning, query processing, text classiï¬cation, hypergraph partitioning.iv
Yazar
Dr. Berkant Barla Cambazoğlu
Bu Yayına Nasıl Atıf Yapılır
Berkant Barla Cambazoğlu (Doctorate thesis). Paralel metin getirme için modeller ve algoritmalar, 2006, Bilkent University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Bilkent University tezlerinden daha fazlası
- Geç Antik Çağ'da Aşağı Tuna: Histria örneği(2023)
- Petrol fiyatları ve getiri eğrisi(2024)
- Sözle yönlendirme üzerine makaleler(2014)
- İletişim ağları ve sağlık uygulamaları için çok kollu haydut algoritmaları(2022)
- Türk Anayasa Mahkemesinin içtihatları ışığında karşılaştırmalı anayasal mutluluk(2023)
- Doğrusal karbon zincirlerinin yoğunluk fonksiyoneli teorisi ile incelenmesi(2023)
