DoctorateOpen Access

Comparison of models estimating overall score and subscore simultaneously in terms of precision, reliability and classification accuracy

2021
0 views
0 downloads
Advisor: Prof. Dr. Hakan Yavuz Atar

Abstract (EN)

This research aims to evaluate the ability estimation models for the overall score and subscores simultaneously. MIRT, HOIRT, and Bifactor models discussed within the scope of the study were compared based on precision, reliability, and classification accuracy. In the study, both simulation data and real data of an English proficiency exam developed and administered by the School of Foreign Languages of a state university in Turkey were used. In the simulation study, the sample size was determined as 5000, the number of items as 30, and the number of dimensions as four. The manipulated variables are the percentage of polytomously-scored items in the total test (5%, 10%, 25%, 50%), test difficulty (very difficult, difficult, medium, easy, very easy), and correlation between dimensions (0.2, 0.5, 0.8). The number of replication was 100 for each of the 60 cross-conditions (3 correlations x 4 levels of percentage of polytomously-scored items x 5 levels of test difficulties), and 6000 data were generated. Overall scores and subscores for MIRT, HOIRT, and Bifactor models were estimated using the BMIRT program. A factorial mixed ANOVA was performed to test the effects of estimation models and simulation conditions on RMSE, reliability, and classification accuracy values of ability estimations. Lastly, the real data were analyzed with all three estimation models, and the standard error, marginal reliability, and classification accuracy values of the ability estimation were examined. In general, the x simulation study results show that the MIRT model outperforms the HOIRT and Bifactor models in all conditions, both in terms of overall score and subscores. When the correlation is high, the difference between the reliability obtained from the estimation models for the overall score is low. For subscores, MIRT and HOIRT have similar results. For overall scores, as the correlation increased, the model performance improved for all three models. For subscores, as the correlation increased, the model performance improved for the MIRT and HOIRT models, while the Bifactor model performance declined. In terms of test difficulty, it was concluded that the models performed better when the test was of medium difficulty, and the highest error, lowest reliability and classification accuracy values were obtained when the test difficulty was very difficult. The highest classification accuracy values were obtained when the test was easy or very easy. There were some differences in the results depending on the levels of the variables, and all of them were reported in detail. Findings obtained with real data analysis also support the simulation study.

Author

Dr. Ayşenur Erdemir

Institution

How to Cite

Ayşenur Erdemir (Doctorate thesis). Comparison of models estimating overall score and subscore simultaneously in terms of precision, reliability and classification accuracy, 2021, Gazi University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Gazi University