Turkish language characteristics and author identification
2009
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Gökhan Dalkılıç
Abstract (EN)
Models of natural languages and language characteristics are widely used in many computer science applications such as data security, language identification, spell checking, data compression, authorship attribution and speech recognition. In the scope of this study, a large scale corpus is created and used to discover language characteristics of Turkish. Word and letter based analyses are made on this corpus to build a base for several NLP studies.In the next step of the study, we used two different methods based on word n-grams to identify author of an anonymous text. For 16 authors, training and test set articles are collected, and mentioned two methods are applied on these article sets. Finally, obtained results are compared and most successful method is determined.
Author
Dr. Feriştah Örücü
Institution
How to Cite
Feriştah Örücü (Master Thesis). Turkish language characteristics and author identification, 2009, Dokuz Eylül University, Bilgisayar Mühendisliği Bölümü.
Keywords
EN
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Dokuz Eylül University
- AFAD gönüllülük sisteminin etkin müdahale açısından analiz(2020)
- Hittite period ceremonial ceramic vessels and current applications(2023)
- The thoughts and practises of Atatürk's adopted daughter Afet İnan(2018)
- Local dependence measures, properties and applications(2007)
- Jurisprudential sanction and influence of Turkish constitutional courts decisions given as a result of norm controlling(2001)
- The Evaluation of treatment costs of childhood cancers(1997)