Master'sOpen Access

Turkish language characteristics and author identification

2009
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Gökhan Dalkılıç

Abstract (EN)

Models of natural languages and language characteristics are widely used in many computer science applications such as data security, language identification, spell checking, data compression, authorship attribution and speech recognition. In the scope of this study, a large scale corpus is created and used to discover language characteristics of Turkish. Word and letter based analyses are made on this corpus to build a base for several NLP studies.In the next step of the study, we used two different methods based on word n-grams to identify author of an anonymous text. For 16 authors, training and test set articles are collected, and mentioned two methods are applied on these article sets. Finally, obtained results are compared and most successful method is determined.

Author

Dr. Feriştah Örücü

How to Cite

Feriştah Örücü (Master Thesis). Turkish language characteristics and author identification, 2009, Dokuz Eylül University, Bilgisayar Mühendisliği Bölümü.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Dokuz Eylül University