Master'sOpen Access

AI-assisted assessment of ESL writing: University instructor perceptions & human-AI scoring correlations

2025
0 views
0 downloads
Advisor: Dr. Öğr. Üyesi Ali İlya

Abstract (EN)

This thesis investigates the integration of artificial intelligence (AI) into English writing assessment through two interconnected studies: a survey of university instructors' perspectives and a correlational analysis of human and AI scoring of argumentative English as a Second Language (ESL) essays. First, a cross-sectional survey (N = 96) captured instructors' usage patterns, trust, perceived effectiveness, ethical concerns, and future outlook regarding AI tools. Usage of general purpose AI tools such as ChatGPT and Grammarly were prevalent. The survey additionally revealed high adoption (82%), moderate confidence in AI feedback for grammar and organization, but considerable reservations about AI's grading capacity, ethical implications, and the necessity of sustained human oversight. Second, a correlational design examined scores assigned by 4 human raters and 25 AI tools on 60 argumentative ESL essays, employing Krippendorff's alpha (Ka), intraclass correlation coefficient (ICC2k), paired t-tests, Wilcoxon signed-rank tests, and Pearson (Pr) and Spearman (Sρ) correlation coefficients. Human raters demonstrated excellent inter-rater reliability (Ka > 0.7; ICC2k > 0.94), while AI tools showed moderate inter-rater reliability (Ka ≤ 0.5; ICC2k > 0.94). AI scores correlated moderately to strongly with human means (Pr ∈ [0.51, 0.65]; Sρ ∈ [0.52, 0.64]) but systematically under-scored by approximately 0.5 to 1 points on all criteria and 2.5 points on total score (p < 0.001; Cohen's d ∈ [-0.98, -0.47]). Among AI tools, OpenAI o3-mini assigned scores which consistently achieved strong correlation with human scores. These findings suggest that, although contemporary AI tools can approximate human judgment in writing assessment, it exhibits systematic bias and limited construct validity. The study underscores the importance of human and AI collaboration, ethical safeguards, and ongoing validation when deploying AI for high-stakes language assessment, pointing to a complementary role for AI that enhances efficiency without displacing human oversight.

Author

Dr. Berkay Yiğit

How to Cite

Berkay Yiğit (Master Thesis). AI-assisted assessment of ESL writing: University instructor perceptions & human-AI scoring correlations, 2025, Sakarya University.

License

Tüm Hakları Saklıdır

This work is shared under the specified license terms.

More theses from Sakarya University