Generalizable Audio Deepfake Detection via Risk-Aware Style Alignment and Structural Empirical Risk Minimization
Mingru Yang, Yanmei Gu, Qianhua He, Peirong Zhang, Haolin He, Zhiming Wang, Huijia Zhu, Jian Liu, Weiqiang Wang
Abstract
With the rapid advancement of AIGC technologies, audio deepfakes have become increasingly realistic, posing serious threats to information security and biometric authentication. Therefore, audio deepfake detection (ADD) has emerged as a critical and fast-evolving research area, particularly requiring superior generalization in out-of-domain scenarios. However, existing ADD methods suffer from constrained generalization and limited access to target data. To address these challenges, we propose Risk-Aware Style Alignment (RASA), a novel generalizable ADD framework that projects the style of any input feature into a shared style space through similarity-based projection. This alignment reduces both inter-domain and intra-source discrepancies without requiring target data during training. In addition, we adopt Structural Empirical Risk Minimization (SERM) in the Poincaré ball model to capture the hierarchical structure of the data and further minimize source risk. By jointly optimizing RASA and SERM, the proposed method effectively tightens the theoretical upper bound of target risk across three key dimensions: source risk, inter-domain divergence, and intra-source discrepancy. Extensive experiments demonstrate that our approach achieves superior generalization and outperforms existing state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 4dbfa192-b5e2-4c09-8d60-1ec4c54e9d15Related papers
- SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake DetectionYi Zhu, Surya Koppisetti, Trang Tran, Gaurav BharajNeurIPS 2024 · 41 citations
- Multi-level SSL Feature Gating for Audio Deepfake DetectionHoan My Tran, Damien Lolive, Aghilas Sini, Arnaud Delhay et al.ACM MM 2025 · 3 citations
- Detect All-Type Deepfake Audio: Wavelet Prompt Tuning for Enhanced Auditory PerceptionYuankun Xie, Ruibo Fu, Xiaopeng Wang, Zhiyong Wang et al.AAAI 2026 · 9 citations
- SafeEar: Content Privacy-Preserving Audio Deepfake DetectionXinfeng Li, Kai Li, Yifan Zheng, Chen Yan et al.CCS 2024 · 26 citations
- Audio-Visual Asynchrony Mitigation: Cross-Modal Alignment and Feature Reconstruction for Deepfake DetectionYan Wang, Qindong Sun, Dongzhu RongACM MM 2025 · 1 citation
