Measuring the Robustness of Audio Deepfake Detectors

Authors:

Xiang Li, Pin-Yu Chen, Wenqi Wei

Where published:

arXiv preprint

2025

 

Abstract:

This paper systematically evaluates how robust audio deepfake detection (ADD) models are under real-world audio distortions. Instead of proposing a new detection method, the authors test 10 different models (including both traditional deep learning models and modern foundation models) against 16 types of audio corruptions, such as noise, audio modifications, and compression. Their experiments reveal that while most models handle noise relatively well, they are highly vulnerable to audio modifications and compression, especially with neural codecs. The study also shows that foundation models outperform traditional models, likely due to large-scale self-supervised pretraining, and that increasing model size improves robustness but with diminishing benefits. Additionally, they demonstrate that targeted data augmentation during training can improve resilience to unseen distortions. A real-world case study on political speech deepfakes further validates their findings. Overall, the paper highlights critical weaknesses in current systems and emphasizes the need for more robust and deployment-ready audio deepfake detection frameworks.

 

Dataset names (used for):

  • Wavefake: main dataset for robustness evaluation.
  • LJSPEECH: source dataset used to derive Wavefake generated audio.
  • In-the-Wild: political speech deepfake case study.
  • ASVSpoof2019 PA: used as reference audio for replay simulation.

 

Some description of the approach:

The paper systematically evaluates robustness of 10 audio deepfake detectors under 16 corruptions: noise perturbation, audio modification, and compression. It compares traditional deep learning models and foundation models, studies model size, tests data augmentation, and includes a political speech case study.

 

Some description of the data:

The study uses the Wavefake dataset, which contains approximately 196 hours of generated audio derived from LJSPEECH. The generated audio was produced using six vocoder architectures: MelGAN, FullBand-MelGAN, MultiBand-MelGAN, HiFi-GAN, Parallel WaveGAN, and WaveGlow.

 

Keywords:

Audio deepfake detection; benchmark; robustness

Instance Represent:

Audio samples / speech waveforms. In preprocessing, audio is resampled to 16 kHz and converted into raw waveforms of 64,000 samples, about 4 seconds.

Dataset Characteristics:

Generated audio deepfake dataset; approximately 196 hours; generated from six vocoder architectures; derived from LJSPEECH; split into train/validation/test. The study also creates corrupted versions using 16 perturbations.

Subject Area:

Audio deepfake detection; AI-generated speech detection; robustness evaluation.

Associated Tools:

Detection models: LFCC-LCNN, ResNet Spec., RawNet2, AASIST, RawGATST, CLAP, Whisper, Wave2Vec2, HuBERT, Wave2Vec2BERT.

Feature Type:

Mel-spectrograms, LFCC, spectrograms, raw waveforms; also foundation model representations.

Number of Instances:

approximately 196 hours of generated audio.

Number of Features:

For LFCC specifically, the paper states 60-dimensional LFCCs are extracted from each audio frame.

Main Paper Link


Last Accessed: 06/16/2026

NSF Award #2346473