Classification of data that are discretized with improved chimerge algorithm with data mining methods
2021
0 görüntülenme
0 i̇ndirme
Danışman: Prof. Dr. Cemalettin Kubat
Özet (EN)
Data discretization can be defined as the process of dividing a continuous feature into a limited number of intervals with the least possible loss of information. Each interval in which the data is divided is assigned a specific value. Discretization is a very important data preprocessing approach for many data mining and machine learning algorithms. Because some algorithms either cannot work with continuous data or show lower performance. On the other hand, discrete data is easier to understand and interpret than continuous data, and it also decreases the running time of different data mining problems such as prediction, classification, and association rules. In this study, four different discretization methods are proposed to increase the performance of the ChiMerge (KiB) algorithm, which is based on the Chi-square statistics and is a widely used discretization method in the literature. The comparative results of these methods, called Elbow ChiMerge (dKiB), Silhouette ChiMerge (sKiB), Square Root ChiMerge (kkKiB), and 2-10 discretization, with the original ChiMerge algorithm, are discussed. Among these methods, dKiB and sKiB are based on finding the most appropriate number of clusters into which the data will be divided by using the k-means algorithm. In the application of these methods, the data set for dKiB is considered as a whole; for sKiB, each attribute of the data set is handled separately. The square root value found for different values of each attribute of the data is determined in the kkKiB algorithm as the number of clusters into which the data will be divided. In 2-10 discretization, the data is divided into 2-10 clusters, respectively, and the results are compared with KiB. Classification success of the methods is measured by stratified 10-fold cross-validation method on 11 real-world datasets using Decision Trees (DT), Naive Bayes (NB), KNearest Neighbors (KNN), and Support Vector Machines (SVM). The obtained results reveal that all four proposed methods generally perform better when compared to the original KiB algorithm.
Yazar
Dr. Nuran Peker
Bu Yayına Nasıl Atıf Yapılır
Nuran Peker (Doctorate thesis). Classification of data that are discretized with improved chimerge algorithm with data mining methods, 2021, Sakarya University.
Anahtar Kelimeler
Lisans
Tüm Hakları Saklıdır
Bu eser belirtilen lisans koşulları altında paylaşılmaktadır.
Sakarya University tezlerinden daha fazlası
- Computational investigation of battery materials using density functional theory(2023)
- Haci Ahmed b. Seyyid al-Bigavî and Tarjama al-Awārif al-maārif (sections of 22-43)(2024)
- Synthesis of carbazol substituted 3,4-dihydropyrimidine-2(1h)-thione deri̇vati̇ves(2024)
- Classification of recyclable wastes with deep learning models: A comparison on the effect of dataset size(2024)
- Hermeneutical analysis of sacrifice, sacred violence and scapegoat motifs in Turkish Mythology(2024)
- Novel thio-chalcone substituted metallophthalocyanines: synthesis, characterization and redox behaviour(2018)
