DoctorateOpen Access

Comparison the effects of missing data imputation methods on classification performance in high dimensional data through simulation

Is this your thesis?

This record came from a bulk archive import. If it’s yours, link it to your profile.

2023
0 views
0 downloads

Abstract (EN)

Objective: This study aims to examine the performance of different missing data imputation methods in accurately estimating missing data in derived high-dimensional datasets and their impact on classification performance using extreme learning machines (ELM). Materials and Methods: In this study, random datasets were generated consisting of n=150 observations with binary dependent variables and p=500 independent variables, considering different data structures, missing data rates, and levels of correlation. Random missing values were created using the missing at random (MAR) mechanism. The missing data imputation methods used in the study included mean, median, random, k-nearest neighbors (KNN), missing value imputation with random forests (I-RF), multiple imputations by chained equations with classification and regression trees (MICE-CART), as well as the direct use of regularized regression (DURR) and the indirect use of regularized regression (IURR) methods developed explicitly for high-dimensional data. Missing values were imputed using these methods. After 1000 iterations of simulations, the performance of the methods in estimating missing values was evaluated based on their proximity of the classification scores obtained using ELM to the reference. Findings: Upon examining the simulation results, according to the applied hierarchical clustering analysis, it was determined that the methods that perform close to each other according to the varying missing rates and correlation levels were in the same cluster. it was observed that in algorithm where variables were associated with a specific set of variables in the dataset, the I-RF, MICE-CART, DURR, IURR, followed by KNN methods exhibited better performance and close to each other and the reference at low missing rates, while the DURR and IURR methods stood out at high missing rates. In the second simulation algorithm, where the data were completely randomly generated, the performances of all methods were found to be close to each other across different correlation levels and missing rates. Conclusion: When the data are completely randomly generated, the prediction performance of the methods used in our study is not affected by the relationships between variables and the missing rates. However, in cases where missing variables are associated with a specific set of variables in the dataset, particularly the DURR and IURR methods prove more effective than the others. These methods were less affected by the relationship between the variables and the variation of the missing rates compared to other methods. Keywords: Extreme Learning Machines, Missing Data, Imputation, Classification, Simulation

Author

Buğra Varol

How to Cite

Buğra Varol (Doctorate thesis). Comparison the effects of missing data imputation methods on classification performance in high dimensional data through simulation, 2023, Aydın Adnan Menderes University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Aydın Adnan Menderes University