Master'sOpen Access

Comparison of techniques and methodologies used in text mining: An application with group meeting speeches of Turkish political part leaders

2011
0 views
0 downloads
Advisor: Doç. Dr. Erman Coşkun

Abstract (EN)

In parallel with the developments in communication and computer technologies, much more information is available today. Collecting information in a very short time, storing, processing, transmitting and transforming it into new information for the demanding departments have given way to the emergence of new disciplines. Text mining is one of these disciplines. Text mining is analyzing un-structured data, namely texts, by means of various statistical methods to extract meaningful and usable information.The first aim of this study is to conduct research on linguistics and technical algorithms which are used in text mining, to compare them and to analyze performance of different classification algorithms with an application. In application part, the aim of this study was to determine by which political party leader the chosen party caucus speeches were made. In this thesis, on this basis, a data set made up of 30 different speeches, every 10 of which were made by one of 3 political leaders, were used. By using parsing method, a feature extraction method, and 2-grams and 3-grams gained from caucus speeches as well as word clustering methods such as K-Means algorithms having characteristics of linguistic and statistical features, 8 different feature vectors were formed. By weighting of these feature vectors were made according to weighting methods of term frequency and term frequency x inverse document frequency. By means of Naive Bayes, a machine learning method, support vector machines, k-nearest neighbor algorithm and decision trees algorithms, the success of each feature vector in classification was compared with that of others.In this study, the most successful classification methods were Naive Bayes and support vector machines. As to classifying documents, 2-grams, gained from caucus speeches, and feature vectors, obtained with the help of K-Means algorithms, were seen to produce more successful results in classifying the speeches.Key words: Text Mining, N-gram, Vector-space Model, Naïve Bayes, Applications of Text Mining

Author

Dr. Keziban Seçkin

How to Cite

Keziban Seçkin (Master Thesis). Comparison of techniques and methodologies used in text mining: An application with group meeting speeches of Turkish political part leaders, 2011, Sakarya University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Sakarya University