Optimizing KNN Classification for Heart Disease Prediction Using Sequential Forward Selection
DOI:
https://doi.org/10.63158/journalisi.v8i4.1744Keywords:
Heart Disease Prediction, Sequential Forward Selection (SFS), K-Nearest Neighbors (KNN), Feature Selection, Wrapper Feature SelectionAbstract
Irrelevant attributes often degrade the effectiveness of distance-based algorithms like K-Nearest Neighbors (KNN) in heart disease prediction. This study enhances a KNN model using Sequential Forward Selection (SFS) on the Cleveland dataset to optimize computational efficiency and precision while maintaining stable recall. To rigorously prevent data leakage, data partitioning was executed prior to mode imputation and normalization, followed by feature selection within a 5-fold stratified cross-validation framework. To further guarantee model robustness and rule out arbitrary selection, a 5-repeated 10-fold cross-validation and a 50-iteration feature stability analysis were executed. Compared to a baseline model (k=7, Euclidean; 80.33% accuracy) utilizing all 13 original attributes, the optimal 7-feature subset (cp, trestbps, chol, thalach, oldpeak, ca, thal) reduced the dimensional space by 46% and achieved 81.97% Accuracy, 77.42% Precision, 85.71% Recall, 78.79% Specificity, 0.9183 ROC-AUC, and an 81.36% F1-Score on an independent hold-out test set comprising 61 samples. Although McNemar's test (p = 1.0000) indicated the absolute accuracy improvement was not statistically significant, the 7-feature model successfully eliminated unstable features and reduced statistical noise. Ultimately, applying SFS provides a highly efficient, computationally lightweight framework for heart disease prediction without compromising predictive reliability.
Downloads
References
[1] F. J. Pinto, E. Harding, and O. Fenton, “A manifesto from Global Heart Hub for early detection and diagnosis of cardiovascular disease,” Eur. Heart J., vol. 46, no. 4, pp. 339–340, Jan. 2025, doi: 10.1093/eurheartj/ehae455.
[2] T. Liu, A. Krentz, L. Lu, and V. Curcin, “Machine learning based prediction models for cardiovascular disease risk using electronic health records data: Systematic review and meta-analysis,” Eur. Heart J. Digit. Health, vol. 6, no. 1, pp. 7–22, Jan. 2025, doi: 10.1093/ehjdh/ztae080.
[3] M. B. Matheson, Y. Kato, S. Baba, C. Cox, J. A. C. Lima, and B. Ambale-Venkatesh, “Cardiovascular risk prediction using machine learning in a large Japanese cohort,” Circ. Rep., vol. 4, no. 12, pp. 595–603, Dec. 2022, doi: 10.1253/circrep.cr-22-0101.
[4] H. J. Kim et al., “Machine learning-based analysis of lifestyle risk factors for atherosclerotic cardiovascular disease: Retrospective case-control study,” JMIR Med. Inform., vol. 13, Art. no. e74415, 2025, doi: 10.2196/74415.
[5] H. Sang et al., “Prediction model for cardiovascular disease in patients with diabetes using machine learning derived and validated in two independent Korean cohorts,” Sci. Rep., vol. 14, no. 1, Dec. 2024, doi: 10.1038/s41598-024-63798-y.
[6] B. Pfeifer and M. Kreuzthaler, “Calibrated kNN classification via second-layer neighborhood analysis,” Adv. Data Anal. Classif., vol. 20, no. 1, pp. 145–165, 2026, doi: 10.1007/s11634-025-00654-5.
[7] D. I. Kasartzian and T. Tsiampalis, “Transforming cardiovascular risk prediction: A review of machine learning and artificial intelligence innovations,” Life, vol. 15, no. 1, Art. no. 94, Jan. 2025, doi: 10.3390/life15010094.
[8] F. Shishehbori and Z. Awan, “Enhancing cardiovascular disease risk prediction with machine learning models,” arXiv Preprint arXiv:2401.17328, Feb. 2024, doi: 10.48550/arXiv.2401.17328.
[9] J. Zeniarja, A. Ukhifahdhina, and A. Salam, “Diagnosis of heart disease using K-nearest neighbor method based on forward selection,” J. Appl. Intell. Syst., vol. 4, no. 2, pp. 39–47, Dec. 2019, doi: 10.33633/jais.v4i2.2749.
[10] Dhiyaussalam, M. H. Noor, Herlinawati, and I. Wardiah, “Impact of feature selection on the performance of KNN and SVM in heart disease prediction,” Tech: J. Eng. Sci., vol. 1, no. 1, pp. 14–25, 2025, doi: 10.69836/tech.v1i1.353.
[11] S. T. Veena, R. Jeevetha, and N. Abirami, “Prediction of heart disease using hybrid feature selection,” J. Posit. Sch. Psychol., vol. 6, no. 4, pp. 10454–10469, 2022.
[12] M. A. Bouqentar et al., “Early heart disease prediction using feature engineering and machine learning algorithms,” Heliyon, vol. 10, no. 19, Oct. 2024, doi: 10.1016/j.heliyon.2024.e38731.
[13] V. B. S. Prasath et al., “Distance and similarity measures effect on the performance of K-nearest neighbor classifier: A review,” arXiv Preprint arXiv:1708.04321, 2017.
[14] C. J. Kelly, A. Karthikesalingam, M. Suleyman, G. Corrado, and D. King, “Key challenges for delivering clinical impact with artificial intelligence,” BMC Med., vol. 17, no. 1, Art. no. 195, Oct. 2019, doi: 10.1186/s12916-019-1426-2.
[15] R. Detrano et al., “Heart Disease,” UCI Mach. Learn. Repository, 1988, doi: 10.24432/C52P4X.
[16] E. N. Wanyonyi and N. W. Masinde, “The impact of data preprocessing on machine learning model performance: A comprehensive examination,” Int. J. Sci. Res. Comput. Sci. Eng. Inf. Technol., vol. 11, no. 2, pp. 3814–3827, Apr. 2025, doi: 10.32628/CSEIT25112854.
[17] S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, no. 9, Sep. 2023, doi: 10.1016/j.patter.2023.100804.
[18] A. Desiani et al., “Handling missing data using combination of deletion technique, mean, mode and artificial neural network imputation for heart disease dataset,” Sci. Technol. Indones., vol. 6, no. 4, pp. 333–342, 2021, doi: 10.26554/sti.2021.6.4.303-312.
[19] A. Alsarhan, F. Hussein, S. Moh, and F. S. El-Salhi, “The effect of preprocessing techniques, applied to numeric features, on classification algorithms’ performance,” Data, vol. 6, no. 7, Art. no. 74, 2021, doi: 10.3390/data6020011.
[20] T. M. Cover and P. E. Hart, “Nearest neighbor pattern classification,” IEEE Trans. Inf. Theory, vol. 13, no. 1, pp. 21–27, Jan. 1967, doi: 10.1109/TIT.1967.1053964.
[21] P. Pudil, J. Novovičová, and J. Kittler, “Floating search methods in feature selection,” Pattern Recognit. Lett., vol. 15, no. 11, pp. 1119–1125, Nov. 1994, doi: 10.1016/0167-8655(94)90127-9.
[22] N. Pudjihartono, T. Fadason, A. W. Kempa-Liehr, and J. M. O’Sullivan, “A review of feature selection methods for machine learning-based disease risk prediction,” Front. Bioinform., vol. 2, Jun. 2022, doi: 10.3389/fbinf.2022.927312.
[23] U. M. Khaire and R. Dhanalakshmi, “Stability of feature selection algorithm: A review,” J. King Saud Univ. Comput. Inf. Sci., vol. 34, no. 4, pp. 1060–1073, Apr. 2022, doi: 10.1016/j.jksuci.2019.06.012.
[24] N. Nasution, M. A. Hasan, and F. B. Nasution, “Predicting heart disease using machine learning: An evaluation of logistic regression, random forest, SVM, and KNN models on the UCI heart disease dataset,” IT J. Res. Dev., vol. 9, no. 2, pp. 140–150, Apr. 2025, doi: 10.25299/itjrd.2025.17941.
[25] R. Mukherjee, S. Sadhu, and A. Kundu, “Heart disease detection using feature selection based KNN classifier,” in Proc. Data Analytics and Management, Singapore: Springer, 2022, doi: 10.1007/978-981-16-6289-8_48.
[26] F. F. Firdaus, H. A. Nugroho, and I. Soesanti, “A review of feature selection and classification approaches for heart disease prediction,” Int. J. Inf. Technol. Electr. Eng., vol. 4, no. 3, Sep. 2020, doi: 10.22146/ijitee.59193.
[27] A. A. Lamir, S. Razzagzadeh, and Z. Rezaei, “A comprehensive machine learning framework for heart disease prediction: Performance evaluation and future perspectives,” arXiv Preprint arXiv:2505.09969, May 2025, doi: 10.48550/arXiv.2505.09969.
[28] H. Takcı, “Performance-enhanced KNN algorithm-based heart disease prediction with the help of optimum parameters,” J. Fac. Eng. Archit. Gazi Univ., vol. 38, no. 1, pp. 451–460, 2023, doi: 10.17341/gazimmfd.977127.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information Systems and Informatics

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors Declaration
- The Authors certify that they have read, understood, and agreed to the Journal of Information Systems and Informatics (JournalISI) submission guidelines, policies, and submission declaration. The submission has been prepared using the provided template.
- The Authors certify that all authors have approved the publication of this manuscript and that there is no conflict of interest.
- The Authors confirm that the manuscript is their original work, has not received prior publication, is not under consideration for publication elsewhere, and has not been previously published.
- The Authors confirm that all authors listed on the title page have contributed significantly to the work, have read the manuscript, attest to the validity and legitimacy of the data and its interpretation, and agree to its submission.
- The Authors confirm that the manuscript is not copied from or plagiarized from any other published work.
- The Authors declare that the manuscript will not be submitted for publication in any other journal or magazine until a decision is made by the journal editors.
- If the manuscript is finally accepted for publication, the Authors confirm that they will either proceed with publication immediately or withdraw the manuscript in accordance with the journal’s withdrawal policies.
- The Authors agree that, upon publication of the manuscript in this journal, they transfer copyright or assign exclusive rights to the publisher, including commercial rights














