An infrastructure model for collecting electronic data to develop large scale corpus
2009
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Yalçın Çebi
Özet (EN)
In the Dokuz Eylül University Computer Engineering Department, different studies on Natural Language Processing (NLP) have been carried out. For NLP research grammatical rules of the language must be determined and a text sample of that language, which is called as corpus, must be prepared. These sample texts should satisfy the grammar rules of language.In this study, an infrastructure for a large scale corpus is designed and implemented. A database model, which supports 6 different document type such as newspaper, report, magazine, book, parliamentary report and official gazette, is designed.By implementing the developed application depending on the database model, 195256 articles were downloaded from 5 newspapers, and their metadata was stored for future use.
Yazar
Dr. Fatma Kızılay
Bu Yayına Nasıl Atıf Yapılır
Fatma Kızılay (Master Thesis). An infrastructure model for collecting electronic data to develop large scale corpus, 2009, Dokuz Eylül University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Dokuz Eylül University tezlerinden daha fazlası
- AFAD gönüllülük sisteminin etkin müdahale açısından analiz(2020)
- Hittite period ceremonial ceramic vessels and current applications(2023)
- Examination of martian habitats from the viewpoint ofstructure(2022)
- Nesnelerin interneti cihazları arasındaki iletişim güvenliğinin arttırılması(2021)
- The thoughts and practises of Atatürk's adopted daughter Afet İnan(2018)
- Sedd ? i Zerai?s being a proof in İslamic Law(2009)
