Comparison of techniques and methodologies used in text mining: An application with group meeting speeches of Turkish political part leaders
2011
0 views
0 downloads
Advisor: Doç. Dr. Erman Coşkun
Abstract (EN)
In parallel with the developments in communication and computer technologies, much more information is available today. Collecting information in a very short time, storing, processing, transmitting and transforming it into new information for the demanding departments have given way to the emergence of new disciplines. Text mining is one of these disciplines. Text mining is analyzing un-structured data, namely texts, by means of various statistical methods to extract meaningful and usable information.The first aim of this study is to conduct research on linguistics and technical algorithms which are used in text mining, to compare them and to analyze performance of different classification algorithms with an application. In application part, the aim of this study was to determine by which political party leader the chosen party caucus speeches were made. In this thesis, on this basis, a data set made up of 30 different speeches, every 10 of which were made by one of 3 political leaders, were used. By using parsing method, a feature extraction method, and 2-grams and 3-grams gained from caucus speeches as well as word clustering methods such as K-Means algorithms having characteristics of linguistic and statistical features, 8 different feature vectors were formed. By weighting of these feature vectors were made according to weighting methods of term frequency and term frequency x inverse document frequency. By means of Naive Bayes, a machine learning method, support vector machines, k-nearest neighbor algorithm and decision trees algorithms, the success of each feature vector in classification was compared with that of others.In this study, the most successful classification methods were Naive Bayes and support vector machines. As to classifying documents, 2-grams, gained from caucus speeches, and feature vectors, obtained with the help of K-Means algorithms, were seen to produce more successful results in classifying the speeches.Key words: Text Mining, N-gram, Vector-space Model, Naïve Bayes, Applications of Text Mining
Author
Dr. Keziban Seçkin
Institution

Sakarya University
Üretim Yönetimi ve Pazarlama Bilim Dalı
How to Cite
Keziban Seçkin (Master Thesis). Comparison of techniques and methodologies used in text mining: An application with group meeting speeches of Turkish political part leaders, 2011, Sakarya University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sakarya University
- Computational investigation of battery materials using density functional theory(2023)
- Haci Ahmed b. Seyyid al-Bigavî and Tarjama al-Awārif al-maārif (sections of 22-43)(2024)
- Synthesis of carbazol substituted 3,4-dihydropyrimidine-2(1h)-thione deri̇vati̇ves(2024)
- Classification of recyclable wastes with deep learning models: A comparison on the effect of dataset size(2024)
- Hermeneutical analysis of sacrifice, sacred violence and scapegoat motifs in Turkish Mythology(2024)
- Novel thio-chalcone substituted metallophthalocyanines: synthesis, characterization and redox behaviour(2018)