Early prediction of student academic performance using deep learning algorithms
2025
0 views
0 downloads
Advisor: Prof. Dr. Orhan Torkul ; Dr. Öğr. Üyesi Tuğba Yıldız
Abstract (EN)
Education is a fundamental process that enhances individuals' quality of life and supports socio-economic development. However, academic failure not only delays students' graduation but also leads to productivity losses at both individual and institutional levels. This situation makes it critically necessary to identify at-risk students at an early stage. Research shows that the overall academic success of a class is most affected by the performance of low-achieving students. Therefore, the progress of these students is considered a key indicator in evaluating the effectiveness of educational processes. Moreover, high failure rates and the resulting need for course repetition negatively impact student motivation and lead to inefficient use of educational resources. For this reason, developing data-driven proactive strategies is of great importance in terms of improving educational quality and ensuring societal development. This study proposes an innovative hybrid model aimed at predicting student success at the beginning of the academic term and identifying individuals in potential risk groups, thereby offering educators the opportunity for early intervention. The model integrates Deep Neural Networks (DNN) with the Particle Swarm Optimization (PSO) algorithm. While DNNs are powerful methods capable of high-accuracy learning in complex and multidimensional datasets, their performance highly depends on effective hyperparameter optimization. At this point, the PSO algorithm serves as an efficient optimization tool and significantly enhances the model's predictive performance. The research was conducted using data from 1,268 students enrolled in the engineering faculty of a public university. The dataset is categorized into three main groups: demographic characteristics, academic performance indicators, and results from high school and higher education entrance exams. A total of 12 predictor variables represents university performance, while 50 variables reflect academic skills acquired during high school. Success status was classified as "pass" or "fail" based on students' ability to pass a specific course. Additionally, the study aims to predict academic success based on high school and entrance exam performance, evaluate the impact of in-term performance data, and identify the most influential attributes. The study was conducted in eight main stages to enable early prediction of students' academic performance. In the first stage, Pearson correlation analysis was applied as a correlation-based feature selection method to eliminate redundant, low-information, or highly interrelated variables from the dataset before model training. Based on a threshold of 0.95, highly correlated variable pairs were identified, and variables with low correlation to GPA were considered unnecessary and removed. Accordingly, "Diploma Grade", "Placement Score", "TYT Score", "SAY Score", "TYT Success Ranking Percentage", and "SAY Success Ranking Percentage" were excluded from the analysis. In the second stage, logistic regression analysis was conducted to assess the impact of independent variables on GPA and to identify the most powerful predictors. Through this analysis, features that significantly increased the explanatory power of the model were evaluated statistically. The variable with the strongest impact on GPA was identified as "Number of Course Repeats", which had a negative effect with a B coefficient of -0.811. Additionally, variables such as "Number of Semesters in Preparatory Program", "Number of Students Taking the Course", "Number of Course Repeats", "High School Achievement Score", and "Physics Correct" were found to be statistically significant at the 5% level. In the third analysis, using the DPH dataset (demographics + academic performance + entrance exam results), the proposed hybrid PSO-DNN model was compared with traditional machine learning and classic DNN methods at the beginning of the term. Experiments showed that the PSO-DNN model achieved the highest accuracy (63.3%), F1 score (56.1%), precision (63.8%), and recall (63.3%). Especially for the "fail" class, the PSO-DNN model yielded higher accuracy and recall. Thus, the PSO-DNN model stands out as a robust and effective approach for early prediction of student performance. In the fourth stage, the effectiveness of different data combinations—used separately and together—on predicting student success at the beginning of the term was examined. The highest accuracy of 65.4% was achieved using the academic (P) data group with the proposed PSO-DNN model. Using the full DPH combination, the model reached 63.3% accuracy, 56.1% F1 score, 63.8% precision, and 63.3% recall. These results indicate that high school grades and entrance exam scores significantly improve prediction accuracy. Overall, the PSO-DNN model outperformed other models in terms of accuracy. The closest results came from the Random Forest model, while Decision Trees showed the lowest performance. These findings highlight that complex model can make better predictions when using more variables. In the fifth analysis, in-term variables such as midterms, quizzes, assignments, and projects were added to the model to evaluate predictions at different times: start of term, before midterms, and before finals. The PSO-DNN model performed best at the start of term and before midterms, while the RF model performed best before finals. Notably, all models showed significant improvements in accuracy, F1 score, precision, and recall before finals, demonstrating that more data improves prediction performance. In the sixth analysis, the model's generalizability was tested using the widely used xAPI-Edu-Data dataset. Despite certain limitations such as swarm size, maximum iterations, and number of hidden layers, the PSO-DNN model achieved performance close to the top two models and better than others in terms of accuracy. In the seventh analysis, explainable artificial intelligence (XAI) techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) were applied to make the model's decision-making process transparent. These analyses identified the most influential features in predictions, including "Number of Students Taking the Course", "High School Achievement Score", "Number of Course Repeats", and "TYT Placement Score". In the final stage, a sensitivity analysis was conducted by comparing XAI findings with t-test and ANOVA results, which showed high consistency, supporting the model's statistical reliability. The analysis revealed that the model is highly sensitive to certain features. Specifically, smaller class sizes had a positive effect on success, axxix balanced course load improved academic performance, preparatory education had statistically significant benefits, and success in STEM subjects contributed significantly to overall academic outcomes. The study's limitations include being limited to one university and one course, technical constraints in hyperparameter optimization, and computational costs. For future work, testing the model in different institutions and disciplines, diversifying optimization techniques, deepening the use of XAI methods, and expanding with cloud-based applications are recommended. Additionally, enhancing data security and privacy will allow for more effective use of the model in practical applications. In conclusion, this study provides significant contributions to data-driven decision-making processes in education. The developed PSO-DNN model has the potential to guide educators and policymakers by predicting students' academic performance at an early stage. Its explainability and transparency can enhance trust in AI-based systems in educational environments. In the future, the widespread adoption of such models may enable the development of personalized education strategies and globally improve student success.
Author
Dr. Ahmet Kala
How to Cite
Ahmet Kala (Doctorate thesis). Early prediction of student academic performance using deep learning algorithms, 2025, Sakarya University.
Keywords
License
Tüm Hakları Saklıdır
This work is shared under the specified license terms.
More theses from Sakarya University
- Computational investigation of battery materials using density functional theory(2023)
- Haci Ahmed b. Seyyid al-Bigavî and Tarjama al-Awārif al-maārif (sections of 22-43)(2024)
- Synthesis of carbazol substituted 3,4-dihydropyrimidine-2(1h)-thione deri̇vati̇ves(2024)
- Classification of recyclable wastes with deep learning models: A comparison on the effect of dataset size(2024)
- Hermeneutical analysis of sacrifice, sacred violence and scapegoat motifs in Turkish Mythology(2024)
- Novel thio-chalcone substituted metallophthalocyanines: synthesis, characterization and redox behaviour(2018)