Edupredict: A Comparative Data Analytics Framework for Student Academic Performance Prediction Using Machine Learning Algorithms

Authors

  • ABRAHAM TEMILADE OLUMIDE JOSEPH AYO BABALOLA UNIVERSITY
  • OLAYINKA OLUSEGUN LAWAL JOSEPH AYO BABALOLA UNIVERSITY
  • TABITHA FUNMILAYO OLUMIDE THE FEDERAL UNIVERSITY OF TECHNOLOGY, AKURE
  • MICHAEL SEGUN OLAJIDE ADEYEMI FEDERAL UNIVERSITY OF EDUCATION, ONDO

Keywords:

, Machine learning, Predictive modelling, Decision Tree, Random Forest, Data Mining

Abstract

The increased demand for data-based decision-making has led to a surge in the use of predictive analytics to identify students at risk of academic failure. However, many tertiary institutions still rely mostly on Cumulative Grade Point Average (CGPA) as an indicator of performance. In this study, we analyzed and compared the predictive capability of three supervised machine learning algorithms (Logistic Regression, Decision Tree, and Random Forest) to forecast student academic outcomes using demographic, behavioral, and academic attributes. We preprocessed the dataset through missing-value handling, feature encoding, and scaling, then split it into 70% training and 30% testing sets. We trained, hyperparameter-tuned, and evaluated the models using accuracy, precision, recall, F1-score, and AUC-ROC, along with Pearson correlation analysis. Results showed that Random Forest performed best (accuracy 85%, AUC-ROC 0.92), outperforming Logistic Regression (78% accuracy, AUC-ROC 0.88) and Decision Tree (74% accuracy, AUC-ROC 0.82). Correlation analysis showed that prior grade was the strongest predictor (r = 0.94), followed by study time, while absences and age showed negligible association; this association is descriptive, not causal. These findings informed the development of Edupredict, a web-based analytics tool that generates real-time probability-of-pass scores from student-level inputs. As a prototype trained entirely on synthetic data, Edupredict demonstrates that such a tool is technically feasible; the reported metrics should be read as evidence that the analytic pipeline works as intended, not as a validated real-world performance, pending future testing on real institutional data.

DOI: https://doi.org/10.5281/zenodo.22981408 

Downloads

Published

2026-09-26