Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake Detection
Yuze Zhao, Kuiyuan Zhang, Zhongyun Hua, Yushu Zhang, Qing Liao, Wei Jiang
Abstract
The rapid evolution of audio deepfakes requires robust detection capable of generalizing to unseen attacks. One-class learning offers inherent robustness for this task by characterizing real speech distributions to detect anomalies. However, establishing a compact decision boundary without spoof supervision remains a fundamental challenge. Existing "relaxed" approaches often compromise this strictness by introducing auxiliary negative samples, which biases the boundary toward seen artifacts and degrades generalization to unseen attacks. To address this, we propose CA-SOADD, a framework that refines the acceptance region by constructing off-manifold boundary probes. Our proposed centroid-anchored tri-objective learning paradigm simultaneously enforces centroid compactness and a centroid-referenced margin against these probes, thereby explicitly tightening the acceptance region without treating them as an explicit negative class. We further extend the framework to heterogeneous settings through domain-conditioned centroids. Experiments on ASVSpoof, CtrSVDD and MLAAD benchmarks demonstrate that our strict real-only method consistently outperforms strong baselines under unseen attack types and domain shifts, with its effectiveness further validated through extensive ablation studies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a19c5670-e10f-4de1-b18b-4e5ade4fff27Builds on5
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 755 citations
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan et al.ICML 2024 · 341 citations
- Improving Generalization for AI-Synthesized Voice DetectionHainan Ren, Li Lin, Chun-Hao Liu, Xin Wang et al.AAAI 2025 · 13 citations
- Negative Data AugmentationAbhishek Sinha, Kumar Ayush, Jiaming Song, Burak Uzkent et al.ICLR 2021 · 3 citations
- CutPaste: Self-Supervised Learning for Anomaly Detection and LocalizationChun-Liang Li, Kihyuk Sohn, Jinsung Yoon, Tomas PfisterCVPR 2021
Related papers
- SONAR: Spectral‑Contrastive Audio Residuals for Generalizable Deepfake DetectionIdo Nitzan Hidekel, Gal Lifshitz, Khen Cohen, Dan RavivICML 2026 · 1 citation
- SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake DetectionYi Zhu, Surya Koppisetti, Trang Tran, Gaurav BharajNeurIPS 2024 · 41 citations
- FakeRadar: Probing Forgery Outliers to Detect Unknown Deepfake VideosZhaolun Li, Jichang Li, Yinqi Cai, Junye Chen et al.ICCV 2025
- Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk MinimizationMingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang et al.ACM MM 2025 · 1 citation
- WhiADD: Semantic-Acoustic Fusion for Robust Audio Deepfake DetectionJianqiao Cui, Bingyao Yu, Qihao Wang, Fei Meng et al.ACM MM 2025 · 1 citation
