Home / Current Issue / Paper 1712477
Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection
Subject area: Science,Engineering and Technology · Area of research: Multilingual Deepfake Audio Detection
Abstract
Deepfake audio refers to synthetic speech that closely mimics a person?s voice, posing risks to security and privacy. This paper proposes a hybrid detection framework combining XLS-R, a multilingual speech representation model, with the Conformer architecture, which captures both local and global audio dependencies. XLS-R extracts rich multilingual embeddings, while the Conformer leverages temporal and contextual features to distinguish genuine from AI-generated speech. Evaluation on benchmark datasets demonstrates that the proposed system achieves improved accuracy and robustness across multiple languages and acoustic conditions.
Keywords
Conformer, Deepfake Audio, Multilingual Speech Representation, XLS-R
References
[1] E. Jain and A. Singh, “Deepfake voice detection using convolutional neural networks: A comprehensive approach to identifying synthetic audio,” in 2024 International Conference on Communication,Control, and Intelligent Systems (CCIS), 2024.
[2] Z. Jin, L. Lang, and B. Leng, “Wave- Spectrogram Cross-Modal Aggregation for Audio Deepfake Detection,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, 10.1109/ICASSP49660.2025.10890563.
[3] R. Mahyavanshi, C. V. Mahesh Reddy, A. J. Shah, and H. A. Patil, “Teager Energy Cepstral Coefficients for Audio Deepfake Detection,” in 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2024, 10.1109/APSIPAASC63619.2025.10848893.
[4] P. Chiddarwar, “Real-Time Detection of AI- Generated Deepfake Audio: A Novel Approach,” in 2024 IEEE 4th International Conference on ICT in Business Industry & Government (ICTBIG), 2024, 10.1109/ICTBIG64922.2024.10911062.
[5] O. A. Shaaban, R. Yildirim, and A. A. Alguttar, “Audio Deepfake Approaches,” IEEE Access, vol. 11, 2023,
[6] M. Gujjar, A. U. Rehman, K. Munir, M. Amjad, and A. Bermak, “Unmasking the Fake: Machine Learning Approach for Deepfake Voice Detection,” IEEE Access, vol. 12, 2024, 10.1109/ACCESS.2024.3521026.
[7] M. Gaikawad and S. Ghosh, “A robust and lightweight CNN-Transformer model for audio deepfake detection in Indian languages,” in 2025 7th International Conference on Signal Processing, Computing and Control (ISPCC), 2025. 10.1109/ISPCC66872.2025.11039572.
[8] T.-P. Doan, L. Nguyen-Vu, S. Jung, and K. Hong, “BTS-E: Audio deepfake detection using breathing-talking-silence encoder,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023. 10.1109/ICASSP49357.2023.10095927.
[9] G. S. Kashyap, S. Kumar, Z. H. Siddiqui, N. Kamuni, M. A. Azeez, J. Gao, and R. Ali, “Fooling the forgers: A multi-stage framework for audio deepfake detection,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025. 10.1109/ICASSP49660.2025.10888175.
[10] D. J. Dsouza, A. P. Rodrigues, and R. Fernandes, “Multi-Modal Comparative Analysis on Audio Dub Detection Using Artificial Intelligence,” IEEE Access, vol 18, 2025,
[11] Y. Xie, H. Cheng, Y. Wang, and L. Ye, “A efficient temporary deepfake location approach based embeddings for partially spoofed audio detection,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024. 10.1109/ICASSP48485.2024.10448196.
[12] D. Song, N. Lee, J. Kim, and E. Choi, “Anomaly detection of deepfake audio based on real audio using generative adversarial network model,” IEEE Access, 2024.
[13] D. U. Leonzio, L. Cuccovillo, P. Bestagini, M. Marcon, P. Aichroth, and S. Tubaro, “Audio Splicing Detection and Localization Based on Acquisition Device Traces,” IEEE Transactions on Information Forensics and Security, vol. 18, 2023.
How to cite this paper
@article{1712477,
author = {Usha Janakiraman, Priyadharshini Ambalavanan, Padmapriya S},
title = {Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {5},
pages = {2438-2446},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1712477.pdf},
abstract = {Deepfake audio refers to synthetic speech that closely mimics a person?s voice, posing risks to security and privacy. This paper proposes a hybrid detection framework combining XLS-R, a multilingual speech representation model, with the Conformer architecture, which captures both local and global audio dependencies. XLS-R extracts rich multilingual embeddings, while the Conformer leverages temporal and contextual features to distinguish genuine from AI-generated speech. Evaluation on benchmark datasets demonstrates that the proposed system achieves improved accuracy and robustness across multiple languages and acoustic conditions.},
keywords = {Conformer, Deepfake Audio, Multilingual Speech Representation, XLS-R},
month = {November},
doi = {https://doi.org/10.64388/IREV9I5-1712477}
}