2026
Robin Singh, Aditya Yogesh Nair, Fabio Palumbo, Florian Barbaro, Anna Dyka, Lohith Rachakonda
The paper conducts a comparative study of audio deepfake detection against different TTS architectures. The authors generate a dataset of 12,000 synthetic speech samples using three modern TTS systems—Dia2, Maya1, and MeloTTS—based on the DailyDialog corpus. They then evaluate four types of detection frameworks that analyze audio from semantic, structural, and signal-level perspectives. Their experiments show that detectors often perform inconsistently: a method that works well for one type of TTS (e.g., streaming or non-autoregressive) may fail on others, especially LLM-based speech generation. To address this, they test a multi-view detection approach that combines multiple analysis levels, which proves to be more robust across all models. Overall, the study demonstrates that single-method detectors are insufficient, and integrating multiple perspectives is essential for reliable audio deepfake detection.
The study uses 4,000 randomly sampled DailyDialog turns to generate 12,000 synthetic audio samples with Dia2, Maya1, and MeloTTS.
Advantages:
Its coverage of newer TTS architectures and its multi-view detection setup, rather than relying on one detection paradigm.
Limitations:
Future work should test adversarial attacks and compression artifacts.
Preprocessing:
The data preprocessing pipeline for audio deepfake detection begins with standardizing all 12,000 audio samples (4,000 each from Dia2, Maya1, and MeloTTS) to a…The data preprocessing pipeline for audio deepfake detection begins with standardizing all 12,000 audio samples (4,000 each from Dia2, Maya1, and MeloTTS) to a uniform 16,000 Hz sample rate, mono channel, and 16-bit PCM format, with each clip trimmed or padded to a fixed duration. Each sample is assigned a binary label (0 = bonafide, 1 = spoof) and optionally sub-labeled by TTS model. Features are then extracted based on the chosen detector — Whisper encoder embeddings for semantic models, multi-layer wav2vec 2.0 representations for XLS-R-SLS, or handcrafted features like MFCC/LFCC for lightweight baselines — followed by mean-variance normalization computed on the training set and applied consistently across all splits. The dataset is divided into 70% training, 15% validation, and 15% test sets using stratified sampling to maintain balanced real-to-fake ratios across all TTS architectures. To improve generalization, augmentation techniques such as additive noise, room impulse response convolution, speed perturbation, and codec compression are applied during training. Finally, quality control metrics including Word Error Rate (WER), Signal-to-Noise Ratio (SNR), Fréchet Audio Distance (FAD), and Speaker Similarity (SIM) are computed to verify the integrity and forensic diversity of the dataset before model training begins.
The paper evaluates three TTS architectures and several detector architectures. Dia2 uses streaming dialogue synthesis with Deep-Inherited Attention; Maya1 uses a 3B Llama backbone and SNAC codec; MeloTTS uses VITS with VAE, flows, and GANs. Detector architectures include Whisper-MesoNet, SSL-AASIST, XLS-R-SLS, MMS-300M, and UncovAI’s proprietary detector.
The evaluation uses OpenAI Whisper-large and the jiwer library for WER calculation, a pre-trained WavLM encoder for speaker similarity, and XLS-R wav2vec 2.0 embeddings for FAD calculation.
Synthetic audio was produced with Dia2, Maya1, and MeloTTS, using DailyDialog text prompts.
Source text dataset: DailyDialog. Generated dataset: 12,000 synthetic audio samples from 4,000 sampled dialogue turns across three TTS systems.
The paper evaluates detection performance using four forensic metrics: EER, AUC, F1-score, and FRR@1%FAR. EER measures the balance between false acceptance and false rejection, AUC reflects overall discrimination ability, F1-score captures precision–recall balance, and FRR@1%FAR evaluates reliability under a strict security threshold.
MeloTTS showed the best generated-audio realism and intelligibility based on FAD and WER, though Maya1 had the best SNR and Dia2 had the highest ACT. Detection performance varied strongly across both detectors and TTS models: Whisper-MesoNet performed best on MeloTTS but worst on Maya1, while XLS-R-SLS and SSL-AASIST performed best on Dia2. UncovAI achieved near-perfect F1 scores across all three synthetic datasets.
The paper contributes a 12,000-sample synthetic dataset, a multi-faceted detection evaluation across four frameworks, evidence that detector robustness differs by TTS architecture, and results showing UncovAI’s proprietary detector achieves near-perfect separation.