Improving Generalization for AI-Synthesized Voice Detection
Hainan Ren, Li Lin, Chun-Hao Liu, Xin Wang, Shu Hu
Abstract
AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 19d154a4-93a9-4b5c-98e4-85eb589b1e2bCited by top-tier papers4
- Rethinking Individual Fairness in Deepfake DetectionAryana Hou, Li Lin, Justin Li, Shu HuACM MM 2025 · 1 citation
- AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness BenchmarkLi Lin, Santosh Santosh, Mingyang Wu, Xin Wang et al.CVPR 2025
- Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake DetectionYuze Zhao, Kuiyuan Zhang, Zhongyun Hua, Yushu Zhang et al.ICML 2026
- MusicDET: Zero-Shot AI-Generated Music DetectionChaolei Han, Hongsong Wang, Jie GuiICML 2026
Builds on12
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
Related papers
- ALDEN: Dual-Level Disentanglement with Meta-learning for Generalizable Audio Deepfake DetectionYuxiong Xu, Bin Li, Weixiang Li, Sara Mandelli et al.ACM MM 2025
- What's the Real: A Novel Design Philosophy for Robust AI-Synthesized Voice DetectionXuan Hai, Xin Liu, Yuan Tan, Gang Liu et al.ACM MM 2024
- Preserving Fairness Generalization in Deepfake DetectionLi Lin, Xinan He, Yan Ju, Xin Wang et al.CVPR 2024
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery DetectionYachao Liang, Min Yu, Gang Li, Jianguo Jiang et al.NeurIPS 2024 · 19 citations
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 6 citations
