Improving Generalization for AI-Synthesized Voice Detection
Hainan Ren, Li Lin, Chun-Hao Liu, Xin Wang, Shu Hu
摘要
AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Rethinking Individual Fairness in Deepfake DetectionAryana Hou, Li Lin, Justin Li, Shu HuACM MM 2025 · 被引用 1 次
- AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness BenchmarkLi Lin, Santosh Santosh, Mingyang Wu, Xin Wang 等CVPR 2025
- Learning Tight Rejection Boundaries without Negatives for Strict One-Class Audio Deepfake DetectionYuze Zhao, Kuiyuan Zhang, Zhongyun Hua, Yushu Zhang 等ICML 2026
- MusicDET: Zero-Shot AI-Generated Music DetectionChaolei Han, Hongsong Wang, Jie GuiICML 2026
它引用的顶会 Paper12
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 被引用 1,267 次
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei 等ICML 2023 · 被引用 773 次
相关 Paper
- ALDEN: Dual-Level Disentanglement with Meta-learning for Generalizable Audio Deepfake DetectionYuxiong Xu, Bin Li, Weixiang Li, Sara Mandelli 等ACM MM 2025
- What's the Real: A Novel Design Philosophy for Robust AI-Synthesized Voice DetectionXuan Hai, Xin Liu, Yuan Tan, Gang Liu 等ACM MM 2024
- Preserving Fairness Generalization in Deepfake DetectionLi Lin, Xinan He, Yan Ju, Xin Wang 等CVPR 2024
- SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery DetectionYachao Liang, Min Yu, Gang Li, Jianguo Jiang 等NeurIPS 2024 · 被引用 19 次
- SiFDetectCracker: An Adversarial Attack Against Fake Voice Detection Based on Speaker-Irrelative FeaturesXuan Hai, Xin Liu, Yuan Tan, Qingguo ZhouACM MM 2023 · 被引用 6 次
