Turkish language characteristics and author identification
2009
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Gökhan Dalkılıç
Abstract (EN)
Models of natural languages and language characteristics are widely used in many computer science applications such as data security, language identification, spell checking, data compression, authorship attribution and speech recognition. In the scope of this study, a large scale corpus is created and used to discover language characteristics of Turkish. Word and letter based analyses are made on this corpus to build a base for several NLP studies.In the next step of the study, we used two different methods based on word n-grams to identify author of an anonymous text. For 16 authors, training and test set articles are collected, and mentioned two methods are applied on these article sets. Finally, obtained results are compared and most successful method is determined.
Author
Dr. Feriştah Örücü
Institution
How to Cite
Feriştah Örücü (Master Thesis). Turkish language characteristics and author identification, 2009, Dokuz Eylül University, Bilgisayar Mühendisliği Bölümü.
Keywords
EN
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Dokuz Eylül University
- Examination of martian habitats from the viewpoint ofstructure(2022)
- Environmental graphic design and public installation in the context of 21st century postmodernism(2022)
- Analysis of speech clarity parameters in open plans offices(2021)
- AFAD gönüllülük sisteminin etkin müdahale açısından analiz(2020)
- The musical analysis of W. A. Mozart, J. N. Hummel and C. M. Von Weber' s bassoon concertos(2006)
- Muscula skeletal injuries of the professional dancers(2006)
