N-gram based classification for turkish text: Author,genre and gender
2006
0 views
0 downloads
Advisor: Y.doç.dr. Banu Diri
Abstract (EN)
We live in a world where the information has an important value. It has been diffucult toaccess the data we need in a reasonable time because of the increasing amount of data,and this has led a new problem of that doing this by hand has been almost impossible.Thus, document classification systems are needed in order to solve the problem. But,there are only a few studies about this topic in Turkish in spite of the other languages.Classification operation is an important subject for document processing, and it allowsdigital documents to be processed automatically. In this thesis study, property vectors ofTurkish language in different dimensions have been constructed by finding out 2nd, 3rd,and 4th order grams. Then, the dimensions of these property vectors have been reducedby using correlation based propoerty choosers, and property vectors in differentdimensions have been obtained. These property vectors have been used in order todetermine the type, author name and author sex of a Turkish document by the help ofclassification methods.The dataset used in the study contains 800 articles of 40 documents which are belong to20 authors writing about sports, magazine, health and politics. The used dataset isarranged in three different formats in order to find out the the type, author name andauthor sex of the documents. Also, 10-times dioganal validity has been applied in order todemonstrate that classification success is not Random.Performances have been compared by using five different types of classification methodsin order to analyze which properties are more successfull for determining type, authorname and author sex of documents. The four known ones of these classification methodsare Naive Bayes, Support Vector Machine, Random Forest, and K-Nearest Neighbormethods. The last of these methods is ng_ind method that we proposed and developed inthis study. Naive Bayes, Support Vector Machine, Random Forest, and K-NearestNeighbor methods have been used together in order to observe the performance of theoperation of using classificators together.The experiments have shown that the articles about sports and daily news are moresuccessfull for determining document types while the articles of women authors are moresuccessfull for determining the author name and author sex of the documents. Theproperty vector obtained by reducing the propoerties has had a better performancecompared to others. DVM method has given the most successful result for authorrecognition while Ng-ind has given the most successful result for genre and genderdetermination. Classificators used together have given more successfull results comparedwith individual classificators.
Author
Sibel Doğan
How to Cite
Sibel Doğan (Master Thesis). N-gram based classification for turkish text: Author,genre and gender, 2006, Yıldız Technical University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Yıldız Technical University
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)
- The annotative dictionary of Kutadgu Bilig in terms of vocabulary(2013)
- Stepper motor speed control with labVIEW(2014)
- Determining supply chain risk factors in food industry(2014)
- TiO2/Cu2O ince film fotovoltaik hücrelerin karakterizasyonu(2014)
