International Peer-Reviewed Journal•Open Access•ISSN 2456-8880
irejournals@gmail.com•+91-7433024337

Home / Current Issue / Paper 1712477

1712477 Vol 9 · Issue 5 Download Paper

Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection

Usha Janakiraman Priyadharshini Ambalavanan Padmapriya S

Subject area: Science,Engineering and Technology  ·  Area of research: Multilingual Deepfake Audio Detection

DOI: 10.64388/IREV9I5-1712477

Abstract

Deepfake audio refers to synthetic speech that closely mimics a person?s voice, posing risks to security and privacy. This paper proposes a hybrid detection framework combining XLS-R, a multilingual speech representation model, with the Conformer architecture, which captures both local and global audio dependencies. XLS-R extracts rich multilingual embeddings, while the Conformer leverages temporal and contextual features to distinguish genuine from AI-generated speech. Evaluation on benchmark datasets demonstrates that the proposed system achieves improved accuracy and robustness across multiple languages and acoustic conditions.

Keywords

Conformer, Deepfake Audio, Multilingual Speech Representation, XLS-R

References

[1] E. Jain and A. Singh, “Deepfake voice detection using convolutional neural networks: A comprehensive approach to identifying synthetic audio,” in 2024 International Conference on Communication,Control, and Intelligent Systems (CCIS), 2024.

[2] Z. Jin, L. Lang, and B. Leng, “Wave- Spectrogram Cross-Modal Aggregation for Audio Deepfake Detection,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025, 10.1109/ICASSP49660.2025.10890563.

[3] R. Mahyavanshi, C. V. Mahesh Reddy, A. J. Shah, and H. A. Patil, “Teager Energy Cepstral Coefficients for Audio Deepfake Detection,” in 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), 2024, 10.1109/APSIPAASC63619.2025.10848893.

[4] P. Chiddarwar, “Real-Time Detection of AI- Generated Deepfake Audio: A Novel Approach,” in 2024 IEEE 4th International Conference on ICT in Business Industry & Government (ICTBIG), 2024, 10.1109/ICTBIG64922.2024.10911062.

[5] O. A. Shaaban, R. Yildirim, and A. A. Alguttar, “Audio Deepfake Approaches,” IEEE Access, vol. 11, 2023,

[6] M. Gujjar, A. U. Rehman, K. Munir, M. Amjad, and A. Bermak, “Unmasking the Fake: Machine Learning Approach for Deepfake Voice Detection,” IEEE Access, vol. 12, 2024, 10.1109/ACCESS.2024.3521026.

[7] M. Gaikawad and S. Ghosh, “A robust and lightweight CNN-Transformer model for audio deepfake detection in Indian languages,” in 2025 7th International Conference on Signal Processing, Computing and Control (ISPCC), 2025. 10.1109/ISPCC66872.2025.11039572.

[8] T.-P. Doan, L. Nguyen-Vu, S. Jung, and K. Hong, “BTS-E: Audio deepfake detection using breathing-talking-silence encoder,” in ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2023. 10.1109/ICASSP49357.2023.10095927.

[9] G. S. Kashyap, S. Kumar, Z. H. Siddiqui, N. Kamuni, M. A. Azeez, J. Gao, and R. Ali, “Fooling the forgers: A multi-stage framework for audio deepfake detection,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2025. 10.1109/ICASSP49660.2025.10888175.

[10] D. J. Dsouza, A. P. Rodrigues, and R. Fernandes, “Multi-Modal Comparative Analysis on Audio Dub Detection Using Artificial Intelligence,” IEEE Access, vol 18, 2025,

[11] Y. Xie, H. Cheng, Y. Wang, and L. Ye, “A efficient temporary deepfake location approach based embeddings for partially spoofed audio detection,” in ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2024. 10.1109/ICASSP48485.2024.10448196.

[12] D. Song, N. Lee, J. Kim, and E. Choi, “Anomaly detection of deepfake audio based on real audio using generative adversarial network model,” IEEE Access, 2024.

[13] D. U. Leonzio, L. Cuccovillo, P. Bestagini, M. Marcon, P. Aichroth, and S. Tubaro, “Audio Splicing Detection and Localization Based on Acquisition Device Traces,” IEEE Transactions on Information Forensics and Security, vol. 18, 2023.

How to cite this paper

Usha Janakiraman, Priyadharshini Ambalavanan, Padmapriya S "Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection" Iconic Research And Engineering Journals Volume 9 Issue 5 2025 Page 2438-2446 https://doi.org/10.64388/IREV9I5-1712477
Usha Janakiraman, Priyadharshini Ambalavanan, Padmapriya S "Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection" Iconic Research And Engineering Journals, vol. 9, no. 5, Nov. 2025, doi: https://doi.org/10.64388/IREV9I5-1712477
Usha Janakiraman, Priyadharshini Ambalavanan, Padmapriya S (2025). Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection. Iconic Research And Engineering Journals, 9(5). doi: https://doi.org/10.64388/IREV9I5-1712477
Usha Janakiraman, Priyadharshini Ambalavanan, Padmapriya S "Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection" Iconic Research And Engineering Journals, vol. 9, no. 5, Nov. 2025. Crossref, https://doi.org/10.64388/IREV9I5-1712477
@article{1712477,
      author = {Usha Janakiraman, Priyadharshini Ambalavanan, Padmapriya S},
      title = {Wav2Vec Meets Conformer: A Novel Hybrid Approach for Multilingual Deepfake Audio Detection},
      journal = {Iconic Research And Engineering Journals},
      year = {2025},
      volume = {9},
      number = {5},
      pages = {2438-2446},
      issn = {2456-8880},
      url = {https://www.irejournals.com/formatedpaper/1712477.pdf},
      abstract = {Deepfake audio refers to synthetic speech that closely mimics a person?s voice, posing risks to security and privacy. This paper proposes a hybrid detection framework combining XLS-R, a multilingual speech representation model, with the Conformer architecture, which captures both local and global audio dependencies. XLS-R extracts rich multilingual embeddings, while the Conformer leverages temporal and contextual features to distinguish genuine from AI-generated speech. Evaluation on benchmark datasets demonstrates that the proposed system achieves improved accuracy and robustness across multiple languages and acoustic conditions.},
      keywords = {Conformer, Deepfake Audio, Multilingual Speech Representation, XLS-R},
      month = {November},
      doi = {https://doi.org/10.64388/IREV9I5-1712477}
  }