Comparison the effects of missing data imputation methods on classification performance in high dimensional data through simulation
Is this your thesis?
This record came from a bulk archive import. If it’s yours, link it to your profile.
Abstract (EN)
Objective: This study aims to examine the performance of different missing data imputation methods in accurately estimating missing data in derived high-dimensional datasets and their impact on classification performance using extreme learning machines (ELM). Materials and Methods: In this study, random datasets were generated consisting of n=150 observations with binary dependent variables and p=500 independent variables, considering different data structures, missing data rates, and levels of correlation. Random missing values were created using the missing at random (MAR) mechanism. The missing data imputation methods used in the study included mean, median, random, k-nearest neighbors (KNN), missing value imputation with random forests (I-RF), multiple imputations by chained equations with classification and regression trees (MICE-CART), as well as the direct use of regularized regression (DURR) and the indirect use of regularized regression (IURR) methods developed explicitly for high-dimensional data. Missing values were imputed using these methods. After 1000 iterations of simulations, the performance of the methods in estimating missing values was evaluated based on their proximity of the classification scores obtained using ELM to the reference. Findings: Upon examining the simulation results, according to the applied hierarchical clustering analysis, it was determined that the methods that perform close to each other according to the varying missing rates and correlation levels were in the same cluster. it was observed that in algorithm where variables were associated with a specific set of variables in the dataset, the I-RF, MICE-CART, DURR, IURR, followed by KNN methods exhibited better performance and close to each other and the reference at low missing rates, while the DURR and IURR methods stood out at high missing rates. In the second simulation algorithm, where the data were completely randomly generated, the performances of all methods were found to be close to each other across different correlation levels and missing rates. Conclusion: When the data are completely randomly generated, the prediction performance of the methods used in our study is not affected by the relationships between variables and the missing rates. However, in cases where missing variables are associated with a specific set of variables in the dataset, particularly the DURR and IURR methods prove more effective than the others. These methods were less affected by the relationship between the variables and the variation of the missing rates compared to other methods. Keywords: Extreme Learning Machines, Missing Data, Imputation, Classification, Simulation
Author
Buğra Varol
How to Cite
Buğra Varol (Doctorate thesis). Comparison the effects of missing data imputation methods on classification performance in high dimensional data through simulation, 2023, Aydın Adnan Menderes University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Aydın Adnan Menderes University
- Determination of some heavy metal levels causing public health risks in honey produced in yatagan province by ICP-MS technique(2021)
- Knowledge and thoughts of women and thei̇r partners related to hysterectomy(2017)
- Appreciation of economic value of natural resources for recreational purposes: a case study on Pamukkale Natural Preservation Area(2018)
- Depression and anxiety level of patients with diabetic foot and associated factors(2018)
- The relationship of chronic idiopathic urticaria with HLA class I and class II antigens(2018)
- Baked clay beak spouted pitchers of II. millennium B.C in Central Anatolia(2006)