Master'sOpen Access

Investigation of text mining methods on Turkish text

2018
0 views
0 downloads
Advisor: Doç. Dr. Sedat Çapar

Abstract (EN)

Today, with the widespread use of the internet the size and value of the data have increased. Making meaningful information from large amounts of data is one step ahead for others and for companies. Various mining techniques must be applied to obtain meaningful information from the data. Data mining processes on structured data. Text mining is the sub-study area of data mining that works on texts. Text mining is the process of analyzing text to extract valuable information from text for special purposes. Before the text mining techniques are applied, the data must be prepared and pre-processed. Extracting meaningful information from texts, classifying text, and reaching the desired information in a short time increase the importance of text mining. Text classification is the process of deciding the class of the given text using the training documents for the predefined categories. The aim in thesis is to classify Turkish data as text. The categories were examined in three categories as "Gender Identification", "Author Identification" and "Species Determination". Naive Bayes method and the bit-score weighting k-NN method were used for classification. The accuracy rates of the two methods are compared. The R programming language is used for classification. In this thesis, a dataset consisting of Turkish columns was created to work on text classification.

Author

Dr. Ezgi Pasin

How to Cite

Ezgi Pasin (Master Thesis). Investigation of text mining methods on Turkish text, 2018, Dokuz Eylül University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Dokuz Eylül University