Hybrid CNN-RNN Architecture with MFCC-LFCC Features for Audio Deepfake Detection

Authors

  • Muh. Hajar Akbar Universitas Sembilanbelas November Kolaka, Indonesia
  • Nurfitria Ningsi Universitas Sembilanbelas November Kolaka, Indonesia
  • Aldi Universitas Sembilanbelas November Kolaka, Indonesia
  • Muhammad Na’im Al Jum’ah Universitas Sembilanbelas November Kolaka, Indonesia
  • Ilcham Universitas Sembilanbelas November Kolaka, Indonesia
Pages Icon

DOI:

https://doi.org/10.63158/journalisi.v8i4.1723

Keywords:

Audio Deepfake Detection, Hybrid CNN-RNN, MFCC, LFCC, Audio Spoofing Detection, Deepfake Audio Forensics

Abstract

The proliferation of sophisticated audio deepfake technology poses a significant threat to digital voice authentication and forensic verification systems. This research addresses this challenge by developing and evaluating a lightweight hybrid Convolutional Neural Network and Recurrent Neural Network (CNN-RNN) architecture for audio deepfake detection. The proposed model integrates a CNN for spatial feature extraction with a bidirectional RNN for temporal dependency modeling, utilizing an early vertically fused feature set of Mel-Frequency Cepstral Coefficients (MFCC) and Linear Frequency Cepstral Coefficients (LFCC) stabilized via utterance-level Cepstral Mean and Variance Normalization (CMVN). We assessed the proposed framework on the official ASVspoof 2019 Logical Access (LA) evaluation benchmark dataset (71,237 trials). Comprehensive evaluation on the official evaluation set demonstrated promising performance within the evaluated benchmark, achieving a Global Equal Error Rate (EER) of 8.24% and an Area Under the Receiver Operating Characteristic (ROC-AUC) of 0.9680, while maintaining an internal validation EER of 0.33% on known attacks. While showing high sensitivity to bona fide speech and robust resilience against advanced Neural Text-to-Speech synthesis (EER < 0.25% for A07–A10), the framework exhibits notable vulnerability to phase-preserving voice conversion attacks. Consequently, without real-world forensic operational testing, the model serves as an initial diagnostic screening approach rather than a fully operational forensic solution, and still requires further cross-repository and noisy-condition validation.

Downloads

Download data is not yet available.

References

[1] N. Chakravarty and M. Dua, "A lightweight feature extraction technique for deepfake audio detection," Multimed. Tools Appl., vol. 83, no. 26, pp. 67443-67467, 2024. doi: 10.1007/s11042-024-18217-9.

[2] H. Radwan, A. Y. Mahmoud, M. A. El-Moneim, and H. M. El-Bakry, "Siam-CNNNet: A novel fusion of Siamese Network and Convolutional Neural Networks based on Mel-Frequency Cepstral Coefficients for audio deepfake detection," in Lecture Notes on Data Engineering and Communications Technologies, vol. 233, Springer, 2024, pp. 193-202. doi: 10.1007/978-3-031-77299-3_19.

[3] K. Duhan and A. Kajal, "Deep learning-based techniques for identification of audio deepfake with open issues: A meta-analysis," SSRG Int. J. Electr. Electron. Eng., vol. 12, no. 3, pp. 36-44, 2025. doi: 10.14445/23488379/IJEEE-V12I3P104.

[4] J. Yi, R. Fu, J. Tao, S. Nie, H. Ma, and C. Wang, "ADD 2022: The first audio deep synthesis detection challenge," in Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP 2022), 2022, pp. 9216-9220. doi: 10.1109/ICASSP43922.2022.9746939.

[5] P. Kawa, M. Plata, and P. Syga, “Attack agnostic dataset: towards generalization and stabilization of audio deepfake detection,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 2022, pp. 4023–4027. doi: 10.21437/Interspeech.2022-10078.

[6] Y. Guo, H. Huang, X. Chen, H. Zhao, and Y. Wang, “Audio deepfake detection with self-supervised wavlm and multi-fusion attentive classifier,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2024, pp. 12702–12706. doi: 10.1109/ICASSP48485.2024.10447923.

[7] J. Khochare, C. Joshi, B. Yenarkar, S. Suratkar, and F. Kazi, “A deep learning framework for audio deepfake detection,” Arab. J. Sci. Eng., vol. 47, no. 3, pp. 3447–3458, 2022, doi: 10.1007/s13369-021-06297-w.

[8] Q. Zhang, S. Wen, and T. Hu, “Audio deepfake detection with self-supervised XLS-R and SLS classifier,” in MM 2024 - Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6765–6773. doi: 10.1145/3664647.3681345.

[9] P. Kawa, M. Plata, and P. Syga, “Defense against adversarial attacks on audio deepfake detection,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 2023, pp. 5276–5280. doi: 10.21437/Interspeech.2023-409.

[10] R. Yan, C. Wen, S. Zhou, T. Guo, W. Zou, and X. Li, “Audio deepfake detection system with neural stitching for add 2022,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2022, pp. 9226–9230. doi: 10.1109/ICASSP43922.2022.9746820.

[11] R. Liu, J. Zhang, and G. Gao, “Multi-space channel representation learning for mono-to-binaural conversion based audio deepfake detection,” Inf. Fusion, vol. 105, 2024, doi: 10.1016/j.inffus.2024.102257.

[12] R. Liu, J. Zhang, G. Gao, and H. Li, “Betray oneself: a novel audio deepfake detection model via mono-to-stereo conversion,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 2023, pp. 3999–4003. doi: 10.21437/Interspeech.2023-2335.

[13] H.-S. Shin, J. Heo, J.-H. Kim, C.-Y. Lim, W. Kim, and H.-J. Yu, “HM-Conformer: A conformer-based audio deepfake detection system with hierarchical pooling and multi-level classification token aggregation methods,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2024, pp. 10581–10585. doi: 10.1109/ICASSP48485.2024.10448453.

[14] V. K. Sharma, R. Garg, and Q. Caudron, “A systematic literature review on deepfake detection techniques,” Multimed. Tools Appl., 2024, doi: 10.1007/s11042-024-19906-1.

[15] I. Amerini et al., “Deepfake media forensics: status and future challenges,” J. Imaging, vol. 11, no. 3, 2025, doi: 10.3390/jimaging11030073.

[16] Z. Khanjani, G. Watson, and V. P. Janeja, “Audio deepfakes: a survey,” Front. Big Data, vol. 5, 2023, doi: 10.3389/fdata.2022.1001063.

[17] J. Li, L. Li, M. Luo, X. Wang, S. Qiao, and Y. Zhou, “Multi-grained backend fusion for manipulation region location of partially fake audio,” in CEUR Workshop Proceedings, 2023, pp. 43–48.

[18] Y. Xu, B. Li, S. Tan, and J. Huang, “Research progress on speech deepfake and its detection techniques,” J. Image Graph., vol. 29, no. 8, pp. 2236–2268, 2024, doi: 10.11834/jig.230476.

[19] Y. Xie et al., “Generalized source tracing: detecting novel audio deepfake algorithm with real emphasis and fake dispersion strategy,” pp. 4833–4837, 2024, doi: 10.21437/interspeech.2024-254.

[20] R. Anagha, A. Arya, V. H. Narayan, S. Abhishek, and T. Anjali, “Audio deepfake detection using deep learning,” in Proceedings of the 2023 12th International Conference on System Modeling and Advancement in Research Trends, SMART 2023, 2023, pp. 176–181. doi: 10.1109/SMART59791.2023.10428163.

[21] J. M. Martín-Doñas and A. Álvarez, “The vicomtech audio deepfake detection system based on wav2vec2 for the 2022 add challenge,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2022, pp. 7937–7941. doi: 10.1109/ICASSP43922.2022.9747768.

[22] A. Khan, K. M. Malik, and S. Nawaz, “Frame-to-utterance convergence: a spectra-temporal approach for unified spoofing detection,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2024, pp. 10761–10765. doi: 10.1109/ICASSP48485.2024.10447500.

[23] J. Frank and L. Schönherr, “WaveFake: A data set to facilitate audio deepfake detection,” Adv. Neural Inf. Process. Syst., no. NeurIPS, 2021.

[24] K. Thakral, H. Agarwal, K. Naraya, S. Mittal, M. Vatsa, and R. Singh, “DeePhyNet: towards detecting phylogeny in deepfakes,” IEEE Trans. Biometrics, Behav. Identity Sci., 2024, doi: 10.1109/TBIOM.2024.3487482.

[25] J. Yang, R. K. Das, and H. Li, “Significance of subband features for synthetic speech detection,” IEEE Trans. Inf. Forensics Secur., vol. 15, pp. 2160–2170, 2020, doi: 10.1109/TIFS.2019.2956589.

[26] Y. Zhang, J. Lu, Z. Li, Z. Shang, W. Wang, and P. Zhang, “Improving the robustness of deepfake audio detection through confidence calibration,” in CEUR Workshop Proceedings, 2023, pp. 70–75.

[27] S. A. Momu, R. R. Siddiqui, S. S. Shanto, and Z. Ahmed, “A comprehensive approach to deepfake audio detection: using feature fusion and deep learning,” in 2024 27th International Conference on Computer and Information Technology, ICCIT 2024 - Proceedings, 2024, pp. 351–356. doi: 10.1109/ICCIT64611.2024.11022599.

[28] K. Schäfer and M. Steinebach, “MFCC vs. LFCC for audio deepfake detection: the role of delta features and input length,” in European Signal Processing Conference, 2025, pp. 576–580. doi: 10.23919/EUSIPCO63237.2025.11226775.

[29] X. Wang et al., “ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech,” 2020, Accessed: Jun. 25, 2026. [Online]. Available: https://www.asvspoof.org/index2019.html

[30] Z. K. Abdul and A. K. Al-Talabani, “Mel frequency cepstral coefficient and its applications: a review,” IEEE Access, vol. 10, pp. 122136–122158, 2022, doi: 10.1109/ACCESS.2022.3223444.

[31] H. Tak, J. Patino, M. Todisco, A. Nautsch, N. Evans, and A. Larcher, “End-to-end anti-spoofing with Rawnet2,” in ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 2021, pp. 6369–6373. doi: 10.1109/ICASSP39728.2021.9414234.

[32] K. Z. Mon, K. Galajit, C. O. Mawalim, J. Karnjana, T. Isshiki, and P. Aimmanee, “Spoof detection using voice contribution on LFCC features and ResNet-34,” in 18th International Conference on Artificial Intelligence and Natural Language Processing and International Conference on Artificial Intelligence and Internet of Things, iSAI-NLP 2023, 2023. doi: 10.1109/iSAI-NLP60301.2023.10354625.

[33] K. S. Krishnan and K. S. Krishnan, “MFAAN: Unveiling audio deepfakes with a multi-feature authenticity network,” in 2023 9th International Conference on Signal Processing and Communication, ICSC 2023, 2023, pp. 150–155. doi: 10.1109/ICSC60394.2023.10441405.

[34] R. Mahyavanshi, C. V Mahesh Reddy, A. J. Shah, and H. A. Patil, “Teager energy cepstral coefficients for audio deepfake detection,” in APSIPA ASC 2024 - Asia Pacific Signal and Information Processing Association Annual Summit and Conference 2024, 2024. doi: 10.1109/APSIPAASC63619.2025.10848893.

[35] J. Zhou and L. Yang, “Research on audio scene classification method based on deep learning technology in sound processing,” in ACM International Conference Proceeding Series, 2024, pp. 664–668. doi: 10.1145/3675417.3675527.

[36] L. Z. Yong and H. Nugroho, “Acoustic anomaly detection of mechanical failure: time-distributed CNN-RNN deep learning models,” in Lecture Notes in Electrical Engineering, 2022, pp. 662–672. doi: 10.1007/978-981-19-3923-5_57.

[37] L. Pham, P. Lam, T. Nguyen, H. Nguyen, and A. Schindler, “Deepfake audio detection using spectrogram-based feature and ensemble of deep learning models,” in IEEE 5th International Symposium on the Internet of Sounds, IS2 2024, 2024. doi: 10.1109/IS262782.2024.10704095.

[38] K. Li, X.-M. Zeng, J.-T. Zhang, and Y. Song, “Convolutional recurrent neural network and multitask learning for manipulation region location,” in CEUR Workshop Proceedings, 2023, pp. 18–22.

[39] H. Tak, J. Patino, A. Nautsch, N. Evans, and M. Todisco, “Spoofing attack detection using the non-linear fusion of sub-band classifiers,” in Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH, 2020, pp. 1106–1110. doi: 10.21437/Interspeech.2020-1844.

[40] R. Lomte, P. Singhal, S. Patil, Z. Shaikh, and P. Agarwal, “A multi-scale residual network with hierarchical feature fusion for robust audio deepfake detection,” in Lecture Notes in Networks and Systems, 2026, pp. 449–466. doi: 10.1007/978-3-032-13803-3_36.

[41] K. Galajit et al., “ThaiSpoof: A database for spoof detection in Thai language,” in 18th International Conference on Artificial Intelligence and Natural Language Processing and International Conference on Artificial Intelligence and Internet of Things, iSAI-NLP 2023, 2023. doi: 10.1109/iSAI-NLP60301.2023.10354956.

Downloads

Published

2026-08-29

Issue

Section

Articles

Most read articles by the same author(s)