When Published:
2025
2025
Authors:
Gueltoum Bendiab; Kamel Zeltni; Mohamed Bader-El-Den; Stavros Shiaeles
Gueltoum Bendiab; Kamel Zeltni; Mohamed Bader-El-Den; Stavros Shiaeles
Abstract:
The paper provides a comprehensive survey of audio deepfake technology by reviewing how modern models generate highly realistic synthetic speech, analyzing both beneficial applications (such as virtual assistants, accessibility, and entertainment) and the associated risks like fraud, misinformation, and impersonation. It examines the evolution of deepfake generation techniques, highlights why these systems are increasingly difficult to detect, and discusses the challenges they pose to existing security mechanisms. Additionally, the authors identify gaps in current detection approaches and outline future research directions aimed at improving robustness, interpretability, and ethical governance in audio deepfake systems.
The paper provides a comprehensive survey of audio deepfake technology by reviewing how modern models generate highly realistic synthetic speech, analyzing both beneficial applications (such as virtual assistants, accessibility, and entertainment) and the associated risks like fraud, misinformation, and impersonation. It examines the evolution of deepfake generation techniques, highlights why these systems are increasingly difficult to detect, and discusses the challenges they pose to existing security mechanisms. Additionally, the authors identify gaps in current detection approaches and outline future research directions aimed at improving robustness, interpretability, and ethical governance in audio deepfake systems.
Evaluation Metrics:
The paper does not conduct experiments, but it compares TTS models using architecture type, synthesis method, speech quality, and inference speed.
The paper does not conduct experiments, but it compares TTS models using architecture type, synthesis method, speech quality, and inference speed.
Contributions:
The paper contributes a systematic review of current audio deepfake generation methods, legitimate applications, misuse risks, cybersecurity challenges, and future research directions.
The paper contributes a systematic review of current audio deepfake generation methods, legitimate applications, misuse risks, cybersecurity challenges, and future research directions.