Tümce kökenli konu modelleme
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Fast augmentation of large text collections in digital world makes inevitable to automatically extract short descriptions of those texts. Even if a lot of studies have been done on detecting hidden topics in text corpora, almost all models follow the bag-of-words assumption. This study presents a new unsupervised learning method that reveals topics in a text corpora and the topic distribution of each text in the corpora. The texts in the corpora are described by a generative graphical model, in which each sentence is generated by a single topic and the topics of consecutive sentences follow a hidden Markov chain. In contrast to bag-of-words paradigm, the model assumes each sentence as a unit block and builds on a memory of topics slowly changing in a meaningful way as the text flows. The results are evaluated both qualitatively by examining topic keywords from particular text collections and quantitatively by means of perplexity, a measure of generalization of the model. Keywords: probabilistic graphical model, topic model, hidden Markov model, Markov chain Monte Carlo.
Author
Can Taylan Sarı
Institution
How to Cite
Can Taylan Sarı (Master Thesis). Tümce kökenli konu modelleme, 2014, İhsan Doğramacı Bilkent University, Bilgisayar Mühendisliği Bölümü.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from İhsan Doğramacı Bilkent University
- Osmanlı Devletinde vergi ve vergi etrafında oluşan ilişkiler üzerine bir çalışma (16.-17. yüzyıllar)(2019)
- Rastsal kümeler ve choquet-tip temsiller(2021)
- Petrol fiyatları ve getiri eğrisi(2024)
- Yalnız yaşamak: Yollar, deneyimler ve gelecek beklentileri(2025)
- Detente dönemine doğru: Johnson Mektubunun ardından Türk dış politikası(2021)
- Geç Antik Çağ'da Aşağı Tuna: Histria örneği(2023)
