Authors:
Mahra Alnaqbi, Richard Adeyemi Ikuesan
Where published:
Discover Computing, 2026, volume 29, article 202.
Published:
2026
Abstract:
This paper conducts a systematic literature review (SLR) of audio deepfake detection research using PRISMA guidelines, analyzing findings from 27 studies. It comprehensively examines detection methods, features, datasets, and evaluation metrics, comparing traditional feature-based approaches (like MFCC and LFCC) with more advanced models such as CNNs, transformers, and multimodal systems. The study finds that while modern deep learning methods achieve higher accuracy, they often fail to generalize well across real-world scenarios and different datasets. It also identifies key limitations in the field, including dataset bias, vulnerability to adversarial attacks, and sensitivity to noise. By synthesizing these insights, the paper highlights a major research gap—the need for a robust and generalizable audio deepfake detection framework that performs consistently across diverse conditions and datasets.
Dataset names (used for):
- The review discusses datasets used for training, testing, and benchmarking audio deepfake detection methods, including DF-TIMIT, DFDC, FakeAVCeleb, WaveFake, ASVspoof 2019, ADD 2022, ADD 2023, Fake-or-Real, CFAD, LAV-DF, In-the-Wild, VoxCeleb2, and others.
Some description of the approach:
The paper uses a PRISMA-compliant systematic literature review approach to analyze audio deepfake detection techniques, features, datasets, metrics, and limitations across selected studies.
Some description of the data:
The review data consists of academic studies selected through database searching, screening, and eligibility assessment. The search retrieved 180 records, screened 130 records, assessed 121 full-text studies, and included 27 studies in the final review.
Keywords:
Audio deepfake, Digital forensics, Machine learning, Multimodal analysis, PRISMA, Systematic literature review
Instance Represent:
N/A
Dataset Characteristics:
The review corpus includes empirical studies on audio deepfake detection published in English.
Subject Area:
Audio deepfake detection, digital forensics, multimedia forensics.
Associated Tools:
The review mentions forensic/detection tools and models such as EnCase, FTK, Wireshark, NetWitness, X1 Social Discovery, and detection models including GMM, SVM, CNN, LCNN, ResNet, RawNet2, SpecRNet, HuBERT, XLS-R, AVTENet, AVTS-DFD, and others.
Feature Type:
The paper reviews cepstral, spectral, time-frequency, temporal, raw-waveform, synchronization, and learned embedding features, including MFCC, LFCC, mel-spectrograms, chromagrams, LFBE, Chroma-STFT, wavelet features, raw audio features, and self-supervised embeddings such as XLS-R and HuBERT.
Number of Instances:
N/A
Number of Features:
N/A