Master'sOpen Access

A multithreaded web crawler and text search engine

2009
0 views
0 downloads
Advisor: Prof. Dr. Selim Akyokuş

Abstract (EN)

Without a doubt, internet is one of the best inventions in the last era. Number of internet users is more than millions. When internet users need information about something or somewhere, they visit search web sites or personal blog pages on the internet. For this purpose, many internet applications have been developed.Search Engines and data mining have shown a big improvement in the last 20 years. The developments on the internet increased the need of accessing and finding correct web resources. Raise of search engines caused to differentiation of search engine services. More intelligent search engines are important for accessing to the correct data.Search engines scan contents of the web sites and create indexes for their contents into own database using robots. Advances in search engines enable classification of subjects of the documents besides words or terms used in a document. Such search engines which have document classification property are called ?Clustered Search Engines?. For determination of page categories, the data mining methods are used.In this thesis study, a web crawler and classification system has been developed. The Open Directory Project (DMOZ) is used as a training set for the classification system. The labeled (categorized) web pages which are stored in the DMOZ directory are used as an input for the classification algorithms. We used classification algorithms available in WEKA Data Mining Tool. The web crawler developed in this thesis classifies web pages according to their subjects while scanning the web pages.

Author

Dr. Arzu Behiye Tarımcı

How to Cite

Arzu Behiye Tarımcı (Master Thesis). A multithreaded web crawler and text search engine, 2009, Doğuş University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Doğuş University