Master'sOpen Access

Evaluation of commonly used missing data methods in terms of descriptive statistics, reliability and validity

2014
0 views
0 downloads
Advisor: Doç. Dr. Zekeriya Nartgün

Abstract (EN)

Missing data is often encountered by researchers. This problem negatively effects the results of researches and causes erroneous inferences. As a solution to this problem different missing data methods were developed. These methods which are used to complete the missing data differ depending on size of sample, quantity of missing data, mechanism of missing data etc.This research is a basic research and in which data sets with different sample size and different amount of missing data were used. The purpose of this study is to define the suitable methods for different conditions by comparing the complete data sets with data sets which are applied 9 different missing data methods, in terms of descriptive statistics, reliability and validity. For this research, Programme for International Students' Assesment (PISA) data were used. PISA 2012 Turkey sample and "Math Work Ethic" scale which was normally distributed and one factored, was selected and data sets which contained 200, 500 and 1000 data were formed at random. Then by using completely missing at random mechanism, %5, %10 and %20 of data were deleted from each data set. In order to complete these missing data, series mean, mean of nearby points, median of nearby points, linear interpolation, linear trend at point, listwise deletion, expectation maksimization, regression imputation and multiple imputation techniques were used. When comparing missing data methods, values obtained from descriptive statistics, reliability and validity were used as referance values. New data values which has been structured with missing data methods compared to referance values in order to make an inference about which method is suitable for different conditions. Some conditiones were compared in descriptive level, some conditiones were compared in terms of t-test and Fisher z test. The results of the study revealed that for different size and missing data rate, listwise deletion method values had the least similarity with the values obtained from complete data sets. Approximate value imputation methods (estimation) values had results which was close to the complete data set values or same with it when the quantity of missing data was low. The methods which had the closest values to values obtained from the complete data set were expectation maximization, regression imputation and multiple imputation. Multiple imputation method outperformed compared to the others. Although, in descriptive comparisons it was found that some methods were more at the forefront, there was no statistically significant difference between methods according to t-test and Fisher z test. Keywords: Missing data methods, descriptive statistics, validity, reliability.

Author

Dr. Merve Şahin Kürşad

How to Cite

Merve Şahin Kürşad (Master Thesis). Evaluation of commonly used missing data methods in terms of descriptive statistics, reliability and validity, 2014, Bolu Abant Izzet Baysal University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Bolu Abant Izzet Baysal University