Comparative Analysis of Four Machine Learning Classifiers for Indonesian Hoax News Detection

Authors

  • Dedi Irawan Universitas Muhammadiyah Metro, Indonesia
  • Sudarmaji Universitas Muhammadiyah Metro, Indonesia
Pages Icon

DOI:

https://doi.org/10.63158/journalisi.v8i4.1710

Keywords:

hoax detection, machine learning, Indonesian language, Random Forest, Bag-of-Words

Abstract

Across Indonesian online platforms, fabricated news spreads faster than fact-checkers can confirm. Because much of the literature relies on resource-intensive deep models, one applied question stays unsettled: which lighter, more transparent classifier best detects Indonesian hoaxes? We assessed four algorithms, Random Forest (RF), Support Vector Machine (SVM), Naïve Bayes (NB), and Extreme Gradient Boosting (XGBoost), on 1,116 Indonesian articles from the MAFINDO/TurnBackHoax repository, using one shared preprocessing pipeline and two feature schemes, Bag-of-Words (BoW) and TF-IDF, under an 80:20 stratified split. On the held-out test set, Random Forest with BoW performed best at 98.66% accuracy and 98.59% macro F1-score, missing three of 224 cases, with XGBoost next at 98.21%. Under repeated five-fold cross-validation, however, XGBoost with BoW attained a significantly higher mean (98.36% versus 97.44% macro F1), so the two ensembles are best regarded as closely competitive rather than decisively ranked. BoW also beat TF-IDF for every classifier, most sharply for Naïve Bayes (84.70% versus 70.19% F1). Because the study uses lexical features alone on a single dataset, the findings indicate that classical ensembles can rival reported deep-learning figures at far lower cost, rather than forming a universal conclusion; broader validation across other sources and periods is still needed.

Downloads

Download data is not yet available.

References

[1] D. K. S. Sekarhati, “Combating Hoax and Misinformation in Indonesia Using Machine Learning: What Is Missing and Future Directions,” Eng. Math. Comput. Sci. J. (EMACS), vol. 6, no. 2, pp. 143–150, 2024. doi: 10.21512/emacsjournal.v6i2.11556

[2] D. Rahmawan, I. Garnesia, and R. Hartanto, “Checking the Fact-Checkers: Analyzing the Content of Fact-Checking Organizations as Initiatives for Hoax Eradication in Indonesia,” J. ASPIKOM, vol. 8, no. 2, pp. 241–256, 2023. doi: 10.24329/aspikom.v8i2.1267

[3] J. Fawaid et al., “TurnBackHoax Dataset: Indonesian News Dataset from turnbackhoax.id,” 2021. [Online]. Available: https://github.com/JibranFawaid/turnbackhoax-dataset.

[4] I. Y. R. Pratiwi, R. A. Asmara, and F. Rahutomo, “Study of hoax news detection using naïve Bayes classifier in Indonesian language,” in Proc. 2017 11th Int. Conf. Inf. Commun. Technol. Syst. (ICTS), Surabaya, Indonesia, 2017, pp. 73–78. doi: 10.1109/ICTS.2017.8265649

[5] F. Rahutomo, I. Y. R. Pratiwi, and D. M. Ramadhani, “Naïve Bayes’s experiment on hoax news detection in Indonesian language,” J. Penelit. Komun. dan Opini Publik, vol. 23, no. 1, pp. 1–15, 2019.

[6] A. Awalina, J. Fawaid, R. Y. Krisnabayu, and N. Yudistira, “Indonesia’s Fake News Detection using Transformer Network,” in Proc. 6th Int. Conf. Sustainable Inf. Eng. Technol. (SIET), Malang, Indonesia: ACM, 2021, pp. 247–251, doi: 10.1145/3479645.3479666.

[7] D. Wijaya, G. M. Sasmitha, and W. O. Vihikan, “Sentiment Analysis of Indonesian Citizens on Electric Vehicle Using FastText and BERT Method,” J. Inf. Syst. Informatics, vol. 6, no. 3, pp. 1360–1372, 2024. doi: 10.51519/journalisi.v6i3.784

[8] O. C. R. Rachmawati and Z. M. E. Darmawan, “The Comparison of Deep Learning Models for Indonesian Political Hoax News Detection,” CommIT J., vol. 18, no. 2, pp. 123–135, 2024.

[9] M. Granik and V. Mesyura, “Fake news detection using naïve Bayes classifier,” in Proc. 2017 IEEE 1st Ukraine Conf. Electr. Comput. Eng. (UKRCON), 2017, pp. 900–903. doi: 10.1109/UKRCON.2017.8100379

[10] F. Panjaitan, W. Ce, H. Oktafiandi, G. Kanugrahan, Y. Ramdhani, and V. H. C. Putra, “Evaluation of Machine Learning Models for Sentiment Analysis in the South Sumatra Governor Election Using Data Balancing Techniques,” J. Inf. Syst. Informatics, vol. 7, no. 1, pp. 461–478, 2025. doi: 10.51519/journalisi.v7i1.1019

[11] P. R. A. Savitri, I M. A. D. Suarjaya, and W. O. Vihikan, “Sentiment Analysis of X (Twitter) Comments on The Influence of South Korean Culture in Indonesia,” J. Inf. Syst. Informatics, vol. 6, no. 2, 2024. doi: 10.51519/journalisi.v6i2.738

[12] A. N. Maulana, W. M. Y. bin Wan Yaacob, and Felawati, “Analyzing Public Sentiment on the Proposal to Return Regional Head Elections to DPRD on Platform X Using the C4.5 Algorithm,” J. Inf. Syst. Informatics, vol. 8, no. 2, pp. 2529–2549, 2026. doi: 10.63158/journalisi.v8i2.1521

[13] F. Sebastiani, “Machine learning in automated text categorization,” ACM Comput. Surv., vol. 34, no. 1, pp. 1–47, 2002. doi: 10.1145/505282.505283

[14] R. N. Devita, H. W. Herwanto, and A. P. Wibawa, “Perbandingan kinerja metode naïve Bayes dan K-nearest neighbor untuk klasifikasi artikel berbahasa Indonesia,” J. Teknol. Inf. dan Ilmu Komput., vol. 5, no. 4, pp. 427–434, 2018. doi: 10.25126/jtiik.201854773

[15] I. Y. R. Pratiwi and A. F. Nugraha, “Hoax news identification using machine learning model from online media in Bahasa Indonesia,” Matrix: J. Manaj. Teknol. dan Inform., vol. 12, no. 2, pp. 58–67, 2022. doi: 10.31940/matrix.v12i2.58-67

[16] I. F. Putra and A. Purwarianti, “Improving Indonesian text classification using multilingual language model,” in Proc. 2020 7th Int. Conf. Adv. Informatics: Concept, Theory Appl. (ICAICTA), Tokoname, Japan: IEEE, 2020, pp. 1–5, doi: 10.1109/ICAICTA51733.2020.9429038.

[17] Y. Setiawan et al., “FakeNews-Mafindo: An Indonesian Fact-Checking Dataset,” 2023. [Online]. Available: https://huggingface.co/datasets/nlp-brin-id/fakenews-mafindo. [Accessed: Jan. 15, 2026].

[18] C. C. Aggarwal and C. Zhai, Mining Text Data. New York: Springer, 2012. doi: 10.1007/978-1-4614-3223-4

[19] L. Breiman, “Random forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001. doi: 10.1023/A:1010933404324

[20] C. Cortes and V. Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995. doi: 10.1007/BF00994018

[21] T. Chen and C. Guestrin, “XGBoost: A scalable tree boosting system,” in Proc. 22nd ACM SIGKDD Int. Conf. Knowl. Discov. Data Min., 2016, pp. 785–794. doi: 10.1145/2939672.2939785

[22] P. Zhang, Y. Jia, and Y. Shang, “Research and application of XGBoost in imbalanced data,” Int. J. Distrib. Sens. Netw., vol. 18, no. 6, 2022. doi: 10.1177/15501329221106935

[23] F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” J. Mach. Learn. Res., vol. 12, pp. 2825–2830, 2011.

[24] O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Sci. Rep., vol. 14, no. 1, p. 6086, 2024. doi: 10.1038/s41598-024-56706-x

[25] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. North Am. Chapter Assoc. Comput. Linguistics (NAACL-HLT), 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423

[26] A. K. Darmawan, M. W. Al Wajieh, M. B. Setyawan, T. Yandi, and H. Hoiriyah, “Hoax News Analysis for the Indonesian National Capital Relocation Public Policy with the Support Vector Machine and Random Forest Algorithms,” J. Inf. Syst. Informatics, vol. 5, no. 1, pp. 150–173, 2023. doi: 10.51519/journalisi.v5i1.438

[27] D. A. N. Krisna and U. Salamah, “Perbandingan algoritma naïve Bayes dan K-nearest neighbor untuk klasifikasi berita hoax kesehatan di media sosial Twitter,” J. Tek. Inform. Kaputama (JTIK), vol. 6, no. 2, 2022.

[28] E. Rasywir and A. Purwarianti, “Eksperimen pada sistem klasifikasi berita hoax berbahasa Indonesia berbasis pembelajaran mesin,” J. Cybermatika, vol. 3, no. 2, 2016.

[29] N. Hoy and T. Koulouri, “An exploration of features to improve the generalisability of fake news detection models,” Expert Syst. Appl., vol. 275, art. 126949, 2025. doi: 10.1016/j.eswa.2025.126949.

Downloads

Published

2026-08-22

Issue

Section

Articles

Most read articles by the same author(s)