Cross-Cohort Early Prediction of Cluster-Derived GPA Trajectories Using First-Year Academic Data and Machine Learning

Authors

DOI:

https://doi.org/10.63158/journalisi.v8i4.1871

Keywords:

Academic risk, Cohort drift, Cross-cohort validation, K-Means trajectory labeling, Student success prediction

Abstract

Early identification of unstable academic development can support timely advising, yet models evaluated only with random splits may not transfer across cohorts. This study examines whether data available after semester-2 grades are finalized can classify two K-Means-derived GPA trajectory labels—stable-improving and fluctuating—observed in semesters 3–7. A retrospective dataset of 1,362 anonymized students was analyzed. Labels were generated from standardized semester 3–7 GPA profiles, while predictors were limited to admission attributes and semester 1–2 records. Train-only preprocessing, RFECV, validation-set threshold tuning, calibration analysis, and GaussianNB, logistic regression, Random Forest, and SVM-RBF were evaluated using a stratified random test and a 2024/2025 administrative-cohort holdout. Random Forest achieved 0.831 balanced accuracy and 0.916 ROC-AUC on the random test. In the holdout, where fluctuating-label prevalence increased from 25.5% to 69.3%, Random Forest had the highest threshold-based balanced accuracy (0.581), whereas SVM-RBF had the highest ROC-AUC (0.763). Drift was substantial in student type, admission semester, status, and study-program composition. In this dataset, random splitting substantially overestimated cross-cohort performance. The observed degradation represents a moderate generalization challenge; these cluster-derived labels require cohort-aware recalibration and educational validation, and predictions should be used only as a human-supervised screening aid.

Downloads

Download data is not yet available.

References

[1] D. Ifenthaler and J. Yau, “Utilising learning analytics to support study success in higher education: a systematic review,” Educ. Technol. Res. Dev., vol. 68, pp. 1961–1990, 2020.

[2] A. Namoun and A. Alshanqiti, “Predicting Student Performance Using Data Mining and Learning Analytics Techniques: A Systematic Literature Review,” Appl. Sci., vol. 11, no. 1, p. 237, 2021.

[3] E. A. Alyahyan and D. Düştegör, “Predicting academic success in higher education: literature review and best practices,” Int. J. Educ. Technol. High. Educ., vol. 17, pp. 1–21, 2020.

[4] N. Sghir, A. Adadi, and M. Lahmer, “Recent advances in Predictive Learning Analytics: A decade systematic review (2012--2022),” Educ. Inf. Technol., vol. 28, pp. 8299–8333, 2023.

[5] Á. Kocsis and G. Molnár, “Factors influencing academic performance and dropout rates in higher education,” Oxford Rev. Educ., vol. 51, pp. 414–432, 2024.

[6] F. Marbouti, J. Ulas, and C. Wang, “Academic and Demographic Cluster Analysis of Engineering Student Success,” IEEE Trans. Educ., vol. 64, pp. 261–266, 2021.

[7] T. Baron et al., “Signatures of medical student applicants and academic success,” PLoS One, vol. 15, 2020.

[8] N. Nawa et al., “Associations between demographic factors and the academic trajectories of medical students in Japan,” PLoS One, vol. 15, no. 5, p. e0233371, 2020.

[9] E. Alhazmi and A. M. Sheneamer, “Early Predicting of Students Performance in Higher Education,” IEEE Access, vol. 11, pp. 27579–27589, 2023.

[10] A. F. Mohamed Nafuri, N. S. Sani, N. F. A. Zainudin, A. H. A. Rahman, and M. Aliff, “Clustering Analysis for Classifying Student Academic Performance in Higher Education,” Appl. Sci., vol. 12, no. 19, p. 9467, 2022.

[11] A. Ridwan, T. Sutikno, I. Riyadi, and W. C. Wahyudin, “On-Time Student Graduation Prediction Modeling: A Comparative Analysis of Naive Bayes Algorithm and Other Data Mining Classifications: Pemodelan Prediksi Kelulusan Mahasiswa Tepat Waktu: Analisis Komparatif Algoritma Naive Bayes Dan Klasifikasi Data Mining Lainnya,” JOINCS (Journal Informatics, Network, Comput. Sci., vol. 8, no. 2, pp. 128–135, 2025.

[12] C. Foster and P. Francis, “A systematic review on the deployment and effectiveness of data analytics in higher education to improve student outcomes,” Assess. Eval. High. Educ., vol. 45, pp. 822–841, 2020.

[13] L. Breiman, “Random Forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.

[14] I. Guyon, J. Weston, S. Barnhill, and V. Vapnik, “Gene selection for cancer classification using support vector machines,” Mach. Learn., vol. 46, no. 1, pp. 389–422, 2002.

[15] C. Cortes and V. Vapnik, “Support-Vector Networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995.

[16] T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006, doi: 10.1016/j.patrec.2005.10.010.

[17] T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLoS One, vol. 10, no. 3, p. e0118432, 2015.

[18] K. H. Brodersen, C. S. Ong, K. E. Stephan, and J. M. Buhmann, “The balanced accuracy and its posterior distribution,” in 20th International Conference on Pattern Recognition, 2010, pp. 3121–3124.

[19] G. Feng, M. Fan, and Y. Chen, “Analysis and Prediction of Students’ Academic Performance Based on Educational Data Mining,” IEEE Access, vol. 10, pp. 19558–19571, 2022.

[20] H. Lim, S. Kim, K.-M. Chung, K. Lee, T. Kim, and J. Heo, “Is college students’ trajectory associated with academic performance?,” Comput. Educ., vol. 178, p. 104397, 2022.

[21] K. M. L. Jones et al., “We’re being tracked at all times: Student perspectives of their privacy in relation to learning analytics in higher education,” J. Assoc. Inf. Sci. Technol., vol. 71, pp. 1044–1059, 2020.

[22] K. M. L. Jones et al., “Transparency and Consent: Student Perspectives on Educational Data Analytics Scenarios,” portal Libr. Acad., vol. 23, pp. 485–515, 2023.

[23] P. Prinsloo, S. Slade, and M. Khalil, “The answer is (not only) technological: Considering student data privacy in learning analytics,” Br. J. Educ. Technol., vol. 53, pp. 876–893, 2022.

[24] C. Fachola, A. Tornaría, P. Bermolen, G. Capdehourat, L. Etcheverry, and M. Fariello, “Federated Learning for Data Analytics in Education,” Data, vol. 8, p. 43, 2023.

[25] W. C. Wahyudin, T. Sutikno, and R. Umar, “A Cluster-Label-Based Framework for Water Quality Risk Pattern Classification Using Naive Bayes and Random Forest,” J. Ilm. Ilmu Terap. Univ. Jambi, vol. 10, no. 3, pp. 1511–1522, 2026, doi: 10.22437/jiituj.v10i3.57126.

[26] M. Yakubu and A. Abubakar, “Applying machine learning approach to predict students’ performance in higher educational institutions,” Kybernetes, vol. 51, no. 2, pp. 916–934, 2021, doi: 10.1108/K-12-2020-0865.

[27] M. V Martins, D. Tolledo, J. Machado, L. M. T. Baptista, and V. Realinho, “Early Prediction of Student’s Performance in Higher Education: A Case Study,” in Trends and Applications in Information Systems and Technologies, WorldCIST 2021, 2021, vol. 1365, pp. 166–175, doi: 10.1007/978-3-030-72657-7_16.

[28] M. Yağcı, “Educational data mining: prediction of students’ academic performance using machine learning algorithms,” Smart Learn. Environ., vol. 9, p. 11, 2022, doi: 10.1186/s40561-022-00192-z.

[29] Y. S. Balcıoğlu and M. Artar, “Predicting academic performance of students with machine learning,” Inf. Dev., vol. 41, no. 3, pp. 896–915, 2025, doi: 10.1177/02666669231213023.

[30] K. Mahawar and P. Rattan, “Empowering education: Harnessing ensemble machine learning approach and ACO-DT classifier for early student academic performance prediction,” Educ. Inf. Technol., vol. 30, pp. 4639–4667, 2025, doi: 10.1007/s10639-024-12976-6.

[31] N. Butt, Z. Mahmood, K. Shakeel, S. Alfarhood, M. S. Safran, and I. Ashraf, “Performance Prediction of Students in Higher Education Using Multi-Model Ensemble Approach,” IEEE Access, vol. 11, pp. 136091–136108, 2023, doi: 10.1109/ACCESS.2023.3336987.

[32] S. Alwarthan, N. Aslam, and I. U. Khan, “Predicting Student Academic Performance at Higher Education Using Data Mining: A Systematic Review,” Appl. Comput. Intell. Soft Comput., vol. 2022, p. 8924028, 2022, doi: 10.1155/2022/8924028.

[33] E. Ahmed, “Student Performance Prediction Using Machine Learning Algorithms,” Appl. Comput. Intell. Soft Comput., vol. 2024, p. 4067721, 2024, doi: 10.1155/2024/4067721.

[34] A. Harif and M. A. Kassimi, “Predictive Modeling of Student Performance Using RFECV-RF for Feature Selection and Machine Learning Techniques,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 7, pp. 228–237, 2024, doi: 10.14569/IJACSA.2024.0150723.

[35] S. Batool, J. Rashid, M. W. Nisar, J. Kim, H.-Y. Kwon, and A. Hussain, “Educational data mining to predict students’ academic performance: A survey study,” Educ. Inf. Technol., vol. 28, pp. 905–971, 2023, doi: 10.1007/s10639-022-11152-y.

[36] E. Tiukhova et al., “Explainable Learning Analytics: Assessing the stability of student success prediction models by means of explainable AI,” Decis. Support Syst., vol. 182, p. 114229, 2024, doi: 10.1016/j.dss.2024.114229.

[37] R. Alamri and B. Alharbi, “Explainable Student Performance Prediction Models: A Systematic Review,” IEEE Access, vol. 9, pp. 33132–33143, 2021, doi: 10.1109/ACCESS.2021.3061368.

[38] F. Arévalo-Cordovilla and M. Peña, “Evaluating ensemble models for fair and interpretable prediction in higher education using multimodal data,” Sci. Rep., vol. 15, art. 29420, 2025, doi: 10.1038/s41598-025-15388-9.

[39] W. C. Wahyudin, T. Sutikno, R. Umar, and A. Ridwan, “Comparison of Data Mining Model Performance in Heart Disease Detection with Feature Selection Application,” JOINCS (Journal Informatics, Network, Comput. Sci., vol. 8, no. 1, pp. 87–93, 2025, doi: 10.21070/joincs.v8i1.1669.

[40] W. C. Wahyudin, T. Sutikno, and R. Umar, “Identification of Bengawan Solo River Water Quality Patterns Using K-Means Clustering Based on Physicochemical and Environmental Parameters,” JOINCS (Journal Informatics, Network, Comput. Sci., vol. 9, no. 1, pp. 43–48, 2026, doi: 10.21070/joincs.v9i1.1710.

[41] L. N. Hakim, F. M. Hana, and W. C. Wahyudin, “Klasifikasi Komentar Toksik Berbahasa Indonesia di Media Sosial Berbasis Fine-Tuning IndoBERT,” JURIKOM (Jurnal Ris. Komputer), vol. 13, no. 1, pp. 202–209, 2026, doi: 10.30865/jurikom.v13i1.9449.

[42] S. P. Afrisia, F. M. Hana, and W. C. Wahyudin, “Implementasi Metode Long Short Term Memory (LSTM) pada Chatbot Kesehatan Mental Mahasiswa,” Sainteks, vol. 21, no. 2, pp. 107–116, 2024, doi: 10.30595/sainteks.v21i2.23869.

Downloads

Published

2026-09-12

Issue

Section

Articles

How to Cite

[1]
A. Ridwan, T. Sutikno, and I. Riadi, “Cross-Cohort Early Prediction of Cluster-Derived GPA Trajectories Using First-Year Academic Data and Machine Learning”, journalisi, vol. 8, no. 4, pp. 5418–5441, Sep. 2026, doi: 10.63158/journalisi.v8i4.1871.