Deepfake Audio Detection in Voice Authentication: A Spectral and CNN-Based Comprehensive Review

Authors:

Ali Osman Mohammed Salih, Abdelmajid Hassan Mansour Emam, Alwalid Bashier Gism Elseed Ahmed, Mahmoud Khalifa, Abdelrazig Suliman, and Nissrein Babiker Mohammed Babiker.

Where published:

Engineering, Technology & Applied Science Research, Vol. 15, No. 6, 2025

 

Abstract:

This study presents a comprehensive review of spectral-based techniques for audio deepfake detection in the context of voice authentication systems. It analyzes how features like spectrograms, MFCC, and CQT can reveal time–frequency inconsistencies in synthetic speech and highlights the effectiveness of integrating CNN-based spoof detection modules as a preprocessing step before identity verification to improve system security. Rather than proposing a new model, the paper synthesizes existing research, identifies key limitations—such as vulnerability to advanced generative models, lack of interpretability, and reduced robustness in noisy real-world conditions—and outlines future directions including hybrid models, adversarial training, and better multilingual datasets. Overall, it provides design insights for building more robust, generalizable, and secure voice authentication systems against deepfake threats.

 

Dataset names (used for):

The review discusses benchmark datasets used in reviewed studies, especially ASVspoof2019, ASVspoof2021, and Fake-or-Real. It also discusses ASVspoof2019 LA and PA protocols for evaluating CNN-based spoof detection systems.

 

Some description of the approach:

The paper is a structured review of spectral and CNN-based audio deepfake detection for voice authentication.

 

Some description of the data:

The “data” for the review are peer-reviewed studies from 2018–2024 found through academic databases.

 

Keywords:

audio deepfakes; voice authentication; spoof detection; spectral features; CNN; ASV spoof.

Dataset Characteristics:

The reviewed datasets are mainly voice spoofing/deepfake benchmarks. ASVspoof2019 includes Logical Access (LA) for TTS/VC synthetic attacks and Physical Access (PA) for replay attacks through microphones/loudspeakers. The review also stresses the need for multilingual, noisy, and real-world datasets.

Subject Area:

Audio deepfake detection, voice authentication, spoof detection, spectral analysis, CNN-based detection, automatic speaker verification security, and audio forensics.

Feature Type:

The main feature types are spectral/time-frequency features, including Mel-spectrogram, CQCC, MFCC, CQT, spectrograms, and raw-waveform features in end-to-end models.

Main Paper Link


License Link


Last Accessed: 06/16/2026

NSF Award #2346473