DoctorateOpen Access

Classification of data that are discretized with improved chimerge algorithm with data mining methods

2021
0 views
0 downloads
Advisor: Prof. Dr. Cemalettin Kubat

Abstract (EN)

Data discretization can be defined as the process of dividing a continuous feature into a limited number of intervals with the least possible loss of information. Each interval in which the data is divided is assigned a specific value. Discretization is a very important data preprocessing approach for many data mining and machine learning algorithms. Because some algorithms either cannot work with continuous data or show lower performance. On the other hand, discrete data is easier to understand and interpret than continuous data, and it also decreases the running time of different data mining problems such as prediction, classification, and association rules. In this study, four different discretization methods are proposed to increase the performance of the ChiMerge (KiB) algorithm, which is based on the Chi-square statistics and is a widely used discretization method in the literature. The comparative results of these methods, called Elbow ChiMerge (dKiB), Silhouette ChiMerge (sKiB), Square Root ChiMerge (kkKiB), and 2-10 discretization, with the original ChiMerge algorithm, are discussed. Among these methods, dKiB and sKiB are based on finding the most appropriate number of clusters into which the data will be divided by using the k-means algorithm. In the application of these methods, the data set for dKiB is considered as a whole; for sKiB, each attribute of the data set is handled separately. The square root value found for different values of each attribute of the data is determined in the kkKiB algorithm as the number of clusters into which the data will be divided. In 2-10 discretization, the data is divided into 2-10 clusters, respectively, and the results are compared with KiB. Classification success of the methods is measured by stratified 10-fold cross-validation method on 11 real-world datasets using Decision Trees (DT), Naive Bayes (NB), KNearest Neighbors (KNN), and Support Vector Machines (SVM). The obtained results reveal that all four proposed methods generally perform better when compared to the original KiB algorithm.

Author

Dr. Nuran Peker

How to Cite

Nuran Peker (Doctorate thesis). Classification of data that are discretized with improved chimerge algorithm with data mining methods, 2021, Sakarya University.

Keywords

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Sakarya University