Edupredict: A Comparative Data Analytics Framework for Student Academic Performance Prediction Using Machine Learning Algorithms
Keywords:
, Machine learning, Predictive modelling, Decision Tree, Random Forest, Data MiningAbstract
The increased demand for data-based decision-making has led to a surge in the use of predictive analytics to identify students at risk of academic failure. However, many tertiary institutions still rely mostly on Cumulative Grade Point Average (CGPA) as an indicator of performance. In this study, we analyzed and compared the predictive capability of three supervised machine learning algorithms (Logistic Regression, Decision Tree, and Random Forest) to forecast student academic outcomes using demographic, behavioral, and academic attributes. We preprocessed the dataset through missing-value handling, feature encoding, and scaling, then split it into 70% training and 30% testing sets. We trained, hyperparameter-tuned, and evaluated the models using accuracy, precision, recall, F1-score, and AUC-ROC, along with Pearson correlation analysis. Results showed that Random Forest performed best (accuracy 85%, AUC-ROC 0.92), outperforming Logistic Regression (78% accuracy, AUC-ROC 0.88) and Decision Tree (74% accuracy, AUC-ROC 0.82). Correlation analysis showed that prior grade was the strongest predictor (r = 0.94), followed by study time, while absences and age showed negligible association; this association is descriptive, not causal. These findings informed the development of Edupredict, a web-based analytics tool that generates real-time probability-of-pass scores from student-level inputs. As a prototype trained entirely on synthetic data, Edupredict demonstrates that such a tool is technically feasible; the reported metrics should be read as evidence that the analytic pipeline works as intended, not as a validated real-world performance, pending future testing on real institutional data.