Audio forensics, Deep learning, Feature extraction, Cross-Channel Augmentation Framework (CCAF) , Spoofing Attacks
AuthorsAbstractSynthetic speech detectors based on deep neural networks (DNN) have achieved close-to-perfect performance in laboratory settings; however, they show poor performance when subject to real-world transmission channels such as Voice over IP (VoIP) and message applications on social media platforms, where lossy compression is common. This is because existing detectors heavily depend on spectral artifacts that are removed by psychoacoustic codecs like Opus and AAC. To overcome this fundamental problem, we introduce a Resilient Forensics architecture based on a Cross-Channel Augmentation Framework (CCAF). In contrast to typical augmentation approaches, CCAF presents a new Codec-Invariant Contrastive Loss that directly reduces the representational gap between highquality audio and their compressed counterparts, hence encouraging channel-invariant feature learning. Our framework uses differentiable codec simulations to incorporate realistic transmission impairments during training. The proposed framework is evaluated on a custom cross-channel partition of the ASVspoof 2021 Logical Access dataset, including a realistic VoIP and streaming compression case study. The Equal Error Rate (EER) is used to measure performance on both clean (matched) and compressed (mismatched) testing data. Our newly proposed CCAF achieves an EER of 5.40% on extreme compression (16 kbps Opus) conditions, which is well below that of the current best-performing baselines including RawNet2 (28.45% EER) and AASIST (21.30% EER). These findings underline the importance of codec invariance in the learning of audio representations for reliable deepfake detection systems in practical scenarios.
•••••••••••••••••••••••••••••••• ejprd.org - Published by Riset Publication Services LLC
EJPRD
Copyright ©2026 by Riset Publication Services LLC