Yüksek LisansAçık Erişim

Author attribution of Turkish documents with hybrid approaches

2006
0 görüntülenme
0 i̇ndirme
Danışman: Yrd. Doç. Dr. Banu Diri

Özet (EN)

There are numerous text documents available in electronic form. With the rapid growthof online information, text categorization has become one of the best automatedtechniques for handling, organizing text data. During the last decades, many classificationtasks that are called author attribution were studied for identifying the author of ananonymous text, or text whose authorship is in doubt.In this study the effect of different features and classifiers on performance of authorattribution of Turkish texts are explored. Different vectors of statistical, grammatical,richness features are generated. Also a set of function words were applied on Turkishdocuments for the first time. All feature sets are combined and new vectors are obtained.In order to escape from features that are not relevant and beneficial for learning, featureselection method is applied over features and new vectors are formed from these reducedfeatures. In the end we obtained 14 different feature vectors.Corpus used in this work is formed from singly-authored 630 documents obtained from35 texts per 18 different authors that are writing on different subjects like medical,popular interest and economics. To determine the capability of identifying authorship forheterogeneous documents, and different dataset sizes, this corpus is divided into 3 parts:Dataset I, Dataset II, Dataset III. Experiments are run 10-fold cross-validation on alldatasets.To analyse which features or feature combinations are successful for identifying theauthor of a document, comparative performance of six different classification methodsare used. These methods are Naive Bayes, Support Vector Machine, Random Forest,Multilayer Perceptron, k-Nearest Neighbour and Self-organizing Feature Vector. Wecombined Random Forest, Naive Bayes and Support Vector Machine in order to analysesuccess ratio in proportion to single classifiers.According to experimental results, most successful results are obtained from corpus ofwhich author count is less and documents are written on different topics. Feature vectorwhich is combined from all features gives better performance than others. Highest scoreis obtained from Multilayer Perceptron method. Combined classifiers gave poor results inproportion to single classifiers.Keywords: Authorship attribution, text classification, feature selection, combiningclassifier, Naive Bayes, Support Vector Machine, Random Forest, Multilayer Perceptron,K-Nearest Neighbour, Self-Organizing Feature Vector.

Yazar

Dr. Filiz Türkoğlu

Bu Yayına Nasıl Atıf Yapılır

Filiz Türkoğlu (Master Thesis). Author attribution of Turkish documents with hybrid approaches, 2006, Yıldız Technical University.

Anahtar Kelimeler

Lisans

Tüm Hakları Saklıdır

Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.

Yıldız Technical University tezlerinden daha fazlası