Comparison of data mining methods in predicting PISA mathematical achievements of students
2018
0 views
0 downloads
Advisor: Prof. Dr. Selahattin Gelbal
Abstract (EN)
The purpose of this study is to examine the performance of Naive Bayes, nearest neighborhood, artificial neural networks, and logistic regression analysis in terms of sample size and test-data ratio in classifying students participated in the PISA (2012) study according to their mathematics performance. The population is students in the 15-year-old group who are participated in the PISA (2012) study. The target population is 62728 students from OECD countries who have participated in the study and have no missing data for the relevant variables. A total of 180 datasets were created by selecting from the target population for the sample sizes including 500 (100 datasets), 1000 (50 datasets) and 5000 (30 datasets) students. The performance of each algorithm was tested by using 11%, 22%, 33%, 44% and 55% of each dataset. It has been checked to what extent the assumptions of the univariate and multivariate analyzes satisfy. For each dataset, 100 analyzes in which test-sample is randomly selected at each time were performed. As the evaluation criteria, accuracy rates and their standard deviations, Kappa values and the area under ROC curve were used. For each dataset, methods' means of accuracy rates and their standard errors were statistically tested. According to the results of the study, while the classification performance of the methods increased as the sample size increased, the increase of the test-data ratio had different effects on the performance of the methods. The Naive Bayes method showed high performance even in small samples, performed the analyzes very quickly and was not affected by the change in the test-data ratio. Logistic regression analysis was the most effective method in large samples, but had poor performance in small samples. While neural networks method showed a similar tendency, its overall performance was lower than Naive Bayes and logistic regression. The lowest performances in all conditions were obtained by the nearest neighbor method. In the conclusions and suggestions part of the present study, the findings were discussed in detail and some suggestions for theory and practice were made.
Author
Dr. İlhan Koyuncu
Institution

Hacettepe University
Eğitimde Ölçme ve Değerlendirme Bilim Dalı
How to Cite
İlhan Koyuncu (Doctorate thesis). Comparison of data mining methods in predicting PISA mathematical achievements of students, 2018, Hacettepe University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Hacettepe University
- Gençlerin ve Gençlik Çalışanlarının Gözünden Gençlik Politikaları ve Hizmetlerinin Değerlendirilmesi(2022)
- Characterization of Ayvalik (Edremit yaglik) extra virgin olive oils volatile compounds with SPME-GC/MS and Raman spectroscopy(2018)
- The effect of child labor related boycott threat on Ivory Coast cocoa production(2018)
- Determination of phonatuary aerodynamic characteristics in turkish speaking children(2018)
- Elementler ve insan doğası arasındaki uyuşmazlık: Ekofobi ve Rönesans İngiliz tiyatrosu(2018)
- Investigation of the presence of carbapenemase in K. pneumoniae and E.coli strains isolated from blood culture by phenotypic and molecular methods(2018)