Yüksek LisansAçık Erişim

Tümce kökenli konu modelleme

2014
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Özgür Ulusoy

Özet (EN)

Fast augmentation of large text collections in digital world makes inevitable to automatically extract short descriptions of those texts. Even if a lot of studies have been done on detecting hidden topics in text corpora, almost all models follow the bag-of-words assumption. This study presents a new unsupervised learning method that reveals topics in a text corpora and the topic distribution of each text in the corpora. The texts in the corpora are described by a generative graphical model, in which each sentence is generated by a single topic and the topics of consecutive sentences follow a hidden Markov chain. In contrast to bag-of-words paradigm, the model assumes each sentence as a unit block and builds on a memory of topics slowly changing in a meaningful way as the text flows. The results are evaluated both qualitatively by examining topic keywords from particular text collections and quantitatively by means of perplexity, a measure of generalization of the model. Keywords: probabilistic graphical model, topic model, hidden Markov model, Markov chain Monte Carlo.

Yazar

Dr. Can Taylan Sarı

Bu Yayına Nasıl Atıf Yapılır

Can Taylan Sarı (Master Thesis). Tümce kökenli konu modelleme, 2014, Bilkent University, Bilgisayar Mühendisliği Bölümü.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Bilkent University tezlerinden daha fazlası