Yüksek LisansAçık Erişim

Investigation of using the LSA model with similarity metrics for semantic-based web document clustering

2018
0 görüntülenme
0 i̇ndirme
Danışman: Yrd. Doç. Dr. Aytuğ Boyacı

Özet (EN)

Web document clustering uses data clustering techniques to group similar web documents into groups, where the documents from the same cluster are more semantically similar than the documents in the other clusters. One of the methods of clustering the documents is based on the topics they contain. The main technique used for topic-based web document clustering is the using of a semantic-analysis model called Latent Semantic Analysis (LSA), which derives a corpus-level semantics (i.e. topics) for every element in the corpus such as, terms and documents. The LSA model has been used in the literature in different ways, variations and for different applications. In this study, we experimentally investigate the best use of the LSA model in semantically clustering the text documents, as there is more than one possible variation when one uses and implements the LSA model. To do so, we examined the LSA model in different combinations with six different semantic-similarity measures to find the best possible variation, which performs best in clustering web documents. The best variation of using the LSA model in text clustering was found after applying it to two commonly used web document datasets. The results also demonstrate the performance of each variation of using LSA model for the task of web document clustering.

Yazar

Dr. Mashhood Alı Alı

Bu Yayına Nasıl Atıf Yapılır

Mashhood Alı Alı (Master Thesis). Investigation of using the LSA model with similarity metrics for semantic-based web document clustering, 2018, Fırat University.

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Fırat University tezlerinden daha fazlası