Author attribution of Turkish documents with hybrid approaches
2006
0 görüntülenme
0 i̇ndirme
Danışman: Yrd. Doç. Dr. Banu Diri
Özet (EN)
There are numerous text documents available in electronic form. With the rapid growthof online information, text categorization has become one of the best automatedtechniques for handling, organizing text data. During the last decades, many classificationtasks that are called author attribution were studied for identifying the author of ananonymous text, or text whose authorship is in doubt.In this study the effect of different features and classifiers on performance of authorattribution of Turkish texts are explored. Different vectors of statistical, grammatical,richness features are generated. Also a set of function words were applied on Turkishdocuments for the first time. All feature sets are combined and new vectors are obtained.In order to escape from features that are not relevant and beneficial for learning, featureselection method is applied over features and new vectors are formed from these reducedfeatures. In the end we obtained 14 different feature vectors.Corpus used in this work is formed from singly-authored 630 documents obtained from35 texts per 18 different authors that are writing on different subjects like medical,popular interest and economics. To determine the capability of identifying authorship forheterogeneous documents, and different dataset sizes, this corpus is divided into 3 parts:Dataset I, Dataset II, Dataset III. Experiments are run 10-fold cross-validation on alldatasets.To analyse which features or feature combinations are successful for identifying theauthor of a document, comparative performance of six different classification methodsare used. These methods are Naive Bayes, Support Vector Machine, Random Forest,Multilayer Perceptron, k-Nearest Neighbour and Self-organizing Feature Vector. Wecombined Random Forest, Naive Bayes and Support Vector Machine in order to analysesuccess ratio in proportion to single classifiers.According to experimental results, most successful results are obtained from corpus ofwhich author count is less and documents are written on different topics. Feature vectorwhich is combined from all features gives better performance than others. Highest scoreis obtained from Multilayer Perceptron method. Combined classifiers gave poor results inproportion to single classifiers.Keywords: Authorship attribution, text classification, feature selection, combiningclassifier, Naive Bayes, Support Vector Machine, Random Forest, Multilayer Perceptron,K-Nearest Neighbour, Self-Organizing Feature Vector.
Yazar
Dr. Filiz Türkoğlu
Bu Yayına Nasıl Atıf Yapılır
Filiz Türkoğlu (Master Thesis). Author attribution of Turkish documents with hybrid approaches, 2006, Yıldız Technical University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Yıldız Technical University tezlerinden daha fazlası
- Examining ?Historical housing structures" within the confines of protecting ecological balance(2012)
- Approximate solutions of integral equations(2012)
- Stepper motor speed control with labVIEW(2014)
- Determining supply chain risk factors in food industry(2014)
- TiO2/Cu2O ince film fotovoltaik hücrelerin karakterizasyonu(2014)
- Study of the problem of evil from a philosophical perspective(2015)
