LHCC: Cross-domain audio deepfake detection via consistency-aware fusion of acoustic and semantic representations

Authors:

Guofu Zhang, Ming Fang, Zhaopin Su, Kejiang Chen, Weiming Zhang, and Yaofei Wang

Where published:

Information Fusion, Volume 134, 2026.

 

Abstract:

The paper proposes a new audio deepfake detection framework (LHCC: Low-High Consistency Checker) designed to improve generalization to unseen deepfake methods. Instead of relying on superficial artifacts, the authors introduce the idea that deepfakes can be detected by inconsistencies between different levels of audio representation. They split a pre-trained model into two parts: a low-level component that captures acoustic details and a high-level component that captures semantic or speaker identity information. They then design an Asymmetric Fusion Module (AFM) to combine these two representations and explicitly measure their consistency, followed by a Dependency Capture Module (DCM) to model temporal patterns of forgery. By focusing on mismatches between low- and high-level features, their method can better detect unseen manipulations. Experiments on multiple cross-domain datasets show that this approach significantly improves generalization and achieves strong performance compared to existing methods.

 

Dataset names (used for):

  • ASVspoof2019 LA: main training dataset and in-domain evaluation.
  • ASVspoof2021 LA and ASVspoof2021 DF: robustness testing under unseen attacks/channel/compression conditions.
  • Fake-or-Real, FakeAVCeleb, and In-the-Wild: cross-dataset/generalization evaluation.

 

Some description of the approach:

The paper proposes LHCC, a cross-domain audio deepfake detector that identifies inconsistencies between low-level acoustic features and high-level semantic/speaker-identity features. It uses XLS-R feature partitioning, AFM for consistency-aware fusion, and DCM to capture temporal forgery patterns, improving generalization to unseen attacks.

 

Some description of the data:

The study uses six benchmark audio deepfake datasets: 19LA, 21LA, 21DF, FoR, FAC, and ITW, covering known and unseen TTS/VC attacks, telephony/channel effects, codec compression, multimodal celebrity deepfakes, and real-world in-the-wild audio.

 

Keywords:

Audio deepfake detection; Information fusion; Consistency-aware fusion; Low-level acoustic representations; Acoustic and semantic representations.

Instance Represent:

Each instance is an audio trial/sample with a bona fide or spoof label. For model input, each audio trial is converted to single-channel 16 kHz and standardized to 64,600 samples by truncation or repetition.

Dataset Characteristics:

The datasets include both bona fide and spoof samples, with official or randomized train/eval partitions.

Subject Area:

Audio deepfake detection, cross-domain generalization, information fusion, self-supervised audio representations, acoustic-semantic consistency, and digital forensics/security.

Associated Tools:

XLS-R encoder, L2-Norm and CKA structural analysis, AFM, DCM, RawBoost augmentation, Adam optimizer, weighted BCE loss, and A100 GPU. Baseline/comparison methods include RawGAT-ST, AASIST, SSLAS, PSDL, AMSDF, SLS, Aletheia, and Nes2Net.

Feature Type:

The main feature type is learning-based SSL representations from XLS-R. LHCC separates XLS-R hidden states into low-level acoustic representations and high-level semantic representations, then fuses them to detect representational inconsistency.

Number of Instances:

947,991

Main Paper Link


Last Accessed: 06/16/2026

NSF Award #2346473