ORPHEUS: A Separation-Robust Proactive Defense for Singing Voice Conversion
Zhaolin Wei, Dengpan Ye, Yanjiao Chen, Jiacheng Deng, Ziyi Liu, Yuhan Lin, Yunna Lv, Yueyun Shang, Zhihong Tian
摘要
Recent advances in singing voice conversion enable realistic cloning of a singer's voice, raising concerns about unauthorized voice misuse. Existing proactive defenses inject imperceptible perturbations to disrupt such conversion, and they are designed for vocal signals rather than mixed music. In real-world scenarios, however, attackers usually obtain mixed music and apply source separation to decompose the mixture into vocal and backing tracks before conversion. Since separation models are trained to distinguish vocals from background components, perturbations are often treated as backing, causing the perturbations to be filtered out. To address this, we present ORPHEUS, the first proactive defense framework tailored for singing voice conversion involving source separation. Specifically, we jointly (i) interfere with the separation model via mask-misguiding and cross-track losses, (ii) disrupt identity information using an ensemble of heterogeneous speaker encoders to enhance transferability, and (iii) optimize perceptual quality through tonality harmony constraints and psychoacoustic masking. To evaluate ORPHEUS, we conduct experiments on two datasets with three source separation models and four singing voice conversion systems. Compared with the best-performing baseline, ORPHEUS further reduces identity similarity by 8.02% and 5.61% on two speaker verification models, respectively, demonstrating consistently stronger defense effectiveness and cross-model transferability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICML 2021 · 被引用 715 次
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen 等AAAI 2022 · 被引用 348 次
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He 等ICLR 2024 · 被引用 75 次
相关 Paper
- SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song CoversGuangke Chen, Yedi Zhang, Fu Song, Ting Wang 等NDSS 2025
- VoiceCloak: A Multi-Dimensional Defense Framework Against Unauthorized Diffusion-Based Voice CloningQianyue Hu, Junyan Wu, Wei Lu, Xiangyang LuoAAAI 2026
- PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial PerturbationsZhaolin Wei, Xiuwen Shi, Dengpan Ye, Yuhan Lin 等ACM MM 2025
- Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice ConversionYanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun 等ACM MM 2023 · 被引用 3 次
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang 等ICML 2025
