SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake Detection
Ido Nitzan Hidekel, Gal Lifshitz, Khen Cohen, Dan Raviv
摘要
Deepfake audio detectors often fail to generalize to unseen attacks, in part due to spectral bias: neural networks prioritize low-frequency structure while under-exploiting subtle high-frequency (HF) artifacts left by generative models. We introduce SONAR (Spectral-cONtrastive Audio Residuals), a frequency-guided framework that explicitly enforces representation-level consistency between semantic content and HF residuals. Unlike prior frequency-aware or dual-stream detectors that treat HF cues as auxiliary features, SONAR encourages structured interaction between content and noise representations in latent space. The model employs a dual-path architecture in which an XLSR encoder captures low-frequency content, while a parallel branch with learnable, value-constrained 1D SRM (Spatial Rich Model) high-pass filters distills HF residuals. The two representations are fused via frequency cross-attention and trained with a Jensen--Shannon alignment loss that promotes LF–HF consistency for genuine audio and amplifies inconsistency for deepfakes. Evaluated on ASVspoof 2021 and in-the-wild benchmarks, SONAR achieves state-of-the-art performance in a single run setting and converges faster than strong baselines. By mitigating the effects of spectral bias through frequency-guided alignment, SONAR provides a fully data-driven and architecture-agnostic approach to generalizable audio deepfake detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain LearningChuangchuang Tan, Yao Zhao, Shikui Wei, Guanghua Gu 等AAAI 2024 · 被引用 232 次
相关 Paper
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency DebiasingHossein Kashiani, Niloufar Alipour Talemi, Fatemeh AfghahCVPR 2025
- On Improving Robustness of Deepfake Image DetectorsAbu Taib Mohammed Shahjahan, Mohammad Mannan, Abdessamad Ben Hamza, Amr YoussefUSENIX Security 2026 · 被引用 1 次
- Audio Deepfake Detection with Self-Supervised XLS-R and SLS ClassifierQishan Zhang, Shuangbing Wen, Tao HuACM MM 2024 · 被引用 54 次
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh 等ICCV 2025 · 被引用 7 次
- Multi-modal Deepfake Detection via Multi-task Audio-Visual Prompt LearningHui Miao, Yuanfang Guo, Zeming Liu, Yunhong WangAAAI 2025 · 被引用 8 次
