Examining computerized adaptive testing under various conditions using different ability estimation methods
2025
0 views
0 downloads
Advisor: Prof. Dr. Bayram Bıçak
Abstract (EN)
This study aims to analyze the performance of computerized adaptive tests (CAT) under different conditions and ability estimation methods. Within the scope of the research, the effects of various ability estimation methods on accuracy, reliability, and efficiency were evaluated using post-hoc simulation results of the CAT system. The findings obtained were used to compare the performance characteristics of ability estimation methods in live CAT applications. Additionally, the study compared CAT and linear tests in terms of various criteria, highlighting the advantages and limitations of both test types. All these applications were designed based on a fundamental research model to test CAT under real and simulation conditions with different ability estimation methods. Within the scope of the research, two different study groups were formed. In the first study group, item test calibration and post-hoc simulations were conducted, while in the second study group, live CAT and linear test applications were carried out. A 155-item pool was developed to measure reading comprehension skills, and data obtained from 1502 middle school students were used to evaluate various starting rules (theta=0, theta=-0.50:0.50), item selection methods (MEI, MFI, KL), ability estimation methods (ML, EAP, MAP), and termination rules (standard error and fixed item number) combinations. Following the post-hoc simulation process, live CAT and a 25-item linear test were administered to 263 students using three different ability estimation methods, and the performance of these methods was compared. The results showed that starting rules had a minimal effect on test performance, except for simulation time and item overlap rates. In terms of item selection methods, MFI was identified as the most efficient method in terms of computation time, while MEI required significantly longer processing times. MEI provided shorter test lengths compared to MFI and KL methods. All item selection methods showed similar correlation values, indicating that they performed similarly in terms of ability estimation accuracy. However, the KL method provided lower RMSE values, offering higher measurement precision. Analyses of termination rules showed that the combination of standard error and fixed item number provided the most accurate results but required longer simulation times. The combination of a 0.50 standard error criterion and 25 items offered an optimal balance between accuracy and efficiency. Among the ability estimation methods, MAP was the fastest method, while ML required the longest computation time. EAP provided the lowest bias and RMSE values, offering higher reliability, especially in short tests. MAP and EAP also showed lower item overlap rates, allowing for more effective use of the item pool. In comparisons between live CAT and linear tests, it was found that CAT, particularly when using MAP and EAP methods, provided more reliable and efficient ability estimations. The ML method produced more errors at extreme ability levels, while the MAP method offered more balanced and stable estimations. Linear tests showed higher variability in ability estimations and required longer test durations. In contrast, CAT, optimized with appropriate starting rules, item selection strategies, and termination criteria, offered significant advantages in terms of accuracy, reliability, and efficiency. Additionally, CAT provided higher measurement information compared to linear tests and performed more precise measurements, especially at medium ability levels. The MAP method provided the highest test information, while the ML method offered relatively lower information levels. Strong correlations were generally observed among live CAT applications, while the relationships between CAT and linear tests were relatively weaker. The ML and EAP methods showed the highest correlations, while the lowest correlation with the linear test was observed with the MAP method, and the highest correlation was observed with the EAP method. These findings indicate that CAT provides more consistent results within itself but may show differences when compared to linear tests. In terms of test duration and item number, the MAP method provided the most efficient results, while the ML method required the longest test duration and the highest number of items. CAT optimized the measurement process by using fewer items and more effectively selecting items appropriate to individuals' ability levels. The fixed item structure of linear tests, which presents the same items to every individual, reduces flexibility in the measurement process, while the adaptive structure of CAT increases measurement accuracy by presenting questions tailored to individuals' ability levels. Keywords: Computerized adaptive testing, item response theory, ability estimation methods, post-hoc simulation, linear test.
Author
Dr. İbrahim Hakkı Tezci
Institution

Akdeniz University
Eğitimde Ölçme ve Değerlendirme Bilim Dalı
How to Cite
İbrahim Hakkı Tezci (Doctorate thesis). Examining computerized adaptive testing under various conditions using different ability estimation methods, 2025, Akdeniz University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Akdeniz University
- Investigation of spin-1 Blume-Capel and mixed spin (1/2, 1) Ising models in the framework of thermodynamic geometry(2024)
- Determining the relationship between air pollution and urbanization and COVID-19 using geographical information systems(2025)
- Identification and mapping of forest fire risk areas; Antalya-Kaş(2025)
- The analysis of values in the works of Christopher Marlowe(2022)
- Andriace Granarium and socio-economic effects(2022)
- Effect of fat, sugar and protein-headed diet on genotoxic potential in Drosophila melanogaster(2022)