Dokument clustering using text mining
2011
0 views
0 downloads
Advisor: Yrd. Doç. Dr. Suat Özdemir
Abstract (EN)
Today, the data in much quantity is kept in type of documents that take place at the internet media. The main problem at here is, to reject the important data from these data and to find out the not discovered patterns. One of the methods that can be used for solving this problem is to find out the relations and patterns between the different groups by grouping of the relations between the documents by using the aggregation techniques. The aggregation analysis has been developed in target of explaining the classification of the objects in details. Related to this target, the elements are separated according to the comparisons inside them. The other target is to make the data set smaller by grouping the alike elements. The target of this study is to prove the necessary data by aggregating the data inside the Turkish and English texts in titles by using the division aggregation techniques. At the study, all texts have been expressed Term Frequency ? Inverse Document Frequency (TF ? IDF) vectors. Later, at the text mining subject, Latin Semantic Index (LSI) method that supplies the deficiency of reaching to the traditional data studies has been used. The LSI method makes up basic concept vectors both from the texts and the terms that are told at these texts by using the K ? Means and K ? Median Algorithms and calculates the projections of each term and text on these vectors. At the study the successes of K ? Means and K ? Median algorithms when TF, TF ? IDF and LSI has been used, has been compared. The aggregating success of K ? Means algorithm has been found better than K ? Median algorithm. At this study, as data set, Milliyet newspaper data set and R8 and WebKB ? 4 data sets that that are frequently used at the literature are used. At Milliyet newspaper data set, there are three subtitles named health, politics and football. R8 data set is found inside Reuters ? 21578 and contents eight classes. WebKB ? 4 data set has been made up by using the web pages that are collected from the computer sciences departments of different universities and contents four classes. The study has been realized by using C# language at Microsoft. Net media.
Author
Dr. Syolai M.taha
How to Cite
Syolai M.taha (Master Thesis). Dokument clustering using text mining, 2011, Gazi University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Gazi University
- Occupational accident analysis and modelling in oil and gas drilling sector Turkey(2021)
- XVI. yüzyıl Anadolu'sunda Oğuzların Karkın Boyu(2004)
- Sharing of real life geometry samples via a social learning environment: A case study(2021)
- Evaluatıon of calcium hydroxide removal efficiency of two different irrigation activation techniques from artificial internal resorption cavities prepared at different root levels(2021)
- Experimental development of the interfacial bond-slip model between textile reinforced mortar strips and masonry walls(2025)
- The use of verbal memory in the context of sustainability and power at the museums of Turk(2010)
