Home / Current Issue / Paper 1711237
Emotional Deepfake Detection Via Voice Stress Analysis
Subject area: Science,Engineering and Technology · Area of research: Artificial Intelligence and Cybersecurity
Abstract
The rapid advancement of generative artificial intelligence has enabled the creation of highly convincing audio deepfakes, where synthetic voices can mimic real speakers with near-human accuracy Posing new threats in fraud, misinformation, and security. Current detection techniques largely rely on acoustic artifacts or signal irregularities, which are increasingly difficult to identify as synthesis models improve. This paper introduces a novel approach for emotional deepfake detection via voice stress analysis. By examining subtle stress and emotion-related cues?such as pitch fluctuations, jitter, shimmer, rhythm, and speech rate? we capture inconsistencies that synthetic voices struggle to replicate. Using emotional speech datasets alongside AI-generated voice samples, we train deep learning models to distinguish authentic from synthetic speech. Results highlight stress-based analysis as a promising defense against evolving deepfake audio attacks
References
[1] Yi et al. (2023) – This comprehensive survey provides an overview of audio deepfake detection, discussing datasets, features, classifiers, and evaluation methods. It highlights the challenges in generalization and the need for interpretability in detection systems.
[2] Zhang et al. (2025) – This paper offers a comprehensive survey of recent advancements in audio deepfake detection, focusing on cutting-edge developments in the past few years..
[3] Li et al. (2025) – The authors propose "Emoanti," a system that utilizes emotion-guided representations for audio anti- deepfake detection. They fine-tune a Wav2Vec2 model on emotion recognition tasks to enhance detection performance.
[4] Wu et al. (2023) – This study addresses individual variabilities in voice stress analysis by incorporating speaker embeddings into hybrid BYOL-S features, significantly improving voice stress detection performance.
[5] Resemble AI (2024) – This article explores how Voice Stress Analysis (VSA) leverages subtle vocal changes to detect deception and assess emotional states, providing insights into the physiological underpinnings of stress-induced speech variations
[6] Mittal et al. (2020) – The authors present a method for detecting deepfake multimedia content by analyzing the similarity between audio and visual modalities and extracting affective cues to infer authenticity.
[7] Aptahire.ai (2025) – Explores how vocal cues such as pitch height and micro tremors increase under emotional tension like guilt or fear during virtual interviews
[8] Warren et al. (2025) – This study demonstrates that prosodic features such as jitter, shimmer, and mean fundamental frequency (F0) can achieve 93% accuracy in detecting audio deepfakes, outperforming traditional artifact-based methods.
[9] Behavioral Signals (2025) – Introduces a real-time deepfake voice detection platform that combines signal analysis with emotion and behavioral intelligence, offering speaker-agnostic protection against synthetic voice threats.
How to cite this paper
@article{1711237,
author = {Asfiya Khanum, Soubiya Siddiqua},
title = {Emotional Deepfake Detection Via Voice Stress Analysis},
journal = {Iconic Research And Engineering Journals},
year = {2025},
volume = {9},
number = {4},
pages = {890-893},
issn = {2456-8880},
url = {https://www.irejournals.com/formatedpaper/1711237.pdf},
abstract = {The rapid advancement of generative artificial intelligence has enabled the creation of highly convincing audio deepfakes, where synthetic voices can mimic real speakers with near-human accuracy Posing new threats in fraud, misinformation, and security. Current detection techniques largely rely on acoustic artifacts or signal irregularities, which are increasingly difficult to identify as synthesis models improve. This paper introduces a novel approach for emotional deepfake detection via voice stress analysis. By examining subtle stress and emotion-related cues?such as pitch fluctuations, jitter, shimmer, rhythm, and speech rate? we capture inconsistencies that synthetic voices struggle to replicate. Using emotional speech datasets alongside AI-generated voice samples, we train deep learning models to distinguish authentic from synthetic speech. Results highlight stress-based analysis as a promising defense against evolving deepfake audio attacks},
month = {October},
}