USENIX Security2026Top-tier venue
ORPHEUS: A Separation-Robust Proactive Defense for Singing Voice Conversion
Zhaolin Wei, Dengpan Ye, Yanjiao Chen, Jiacheng Deng, Ziyi Liu, Yuhan Lin, Yunna Lv, Yueyun Shang, Zhihong Tian
Abstract
Recent advances in singing voice conversion enable realistic cloning of a singer's voice, raising concerns about unauthorized voice misuse. Existing proactive defenses inject imperceptible perturbations to disrupt such conversion, and they are designed for vocal signals rather than mixed music. In real-world scenarios, however, attackers usually obtain mixed music and apply source separation to decompose the mixture into vocal and backing tracks before conversion. Since separation models are trained to distinguish vocals from background components, perturbations are often treated as backing, causing the perturbations to be filtered out. To address this, we present ORPHEUS, the first proactive defense framework tailored for singing voice conversion involving source separation. Specifically, we jointly (i) interfere with the separation model via mask-misguiding and cross-track losses, (ii) disrupt identity information using an ensemble of heterogeneous speaker encoders to enhance transferability, and (iii) optimize perceptual quality through tonality harmony constraints and psychoacoustic masking. To evaluate ORPHEUS, we conduct experiments on two datasets with three source separation models and four singing voice conversion systems. Compared with the best-performing baseline, ORPHEUS further reduces identity similarity by 8.02% and 5.61% on two speaker verification models, respectively, demonstrating consistently stronger defense effectiveness and cross-model transferability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d584da8-6b11-43fb-bdbf-fb01f434a054Builds on12
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- Grad-TTS: A Diffusion Probabilistic Model for Text-to-SpeechVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova et al.ICML 2021 · 715 citations
- DiffSinger: Singing Voice Synthesis via Shallow Diffusion MechanismJinglin Liu, Chengxi Li, Yi Ren, Feiyang Chen et al.AAAI 2022 · 348 citations
- Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech SynthesisZiyue Jiang, Jinglin Liu, Yi Ren, Jinzheng He et al.ICLR 2024 · 75 citations
Related papers
- SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song CoversGuangke Chen, Yedi Zhang, Fu Song, Ting Wang et al.NDSS 2025
- VoiceCloak: A Multi-Dimensional Defense Framework Against Unauthorized Diffusion-Based Voice CloningQianyue Hu, Junyan Wu, Wei Lu, Xiangyang LuoAAAI 2026
- PhonoFence: A Cross-Task Defense Framework for DeepFake via Phoneme-Level Adversarial PerturbationsZhaolin Wei, Xiuwen Shi, Dengpan Ye, Yuhan Lin et al.ACM MM 2025
- Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice ConversionYanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun et al.ACM MM 2023 · 3 citations
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang et al.ICML 2025
