Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice Conversion
Yanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun, Rubing Shen, Lina Wang
Abstract
Voice conversion (VC), as a voice style transfer technology, is becoming increasingly prevalent while raising serious concerns about its illegal use. Proactively tracing the origins of VC-generated speeches, i.e., speaker traceability, can prevent the misuse of VC, but unfortunately has not been extensively studied. In this paper, we are the first to investigate the speaker traceability for VC and propose a traceable VC framework named VoxTracer. Our VoxTracer is similar to but beyond the paradigm of audio watermarking. We first use unique speaker embedding to represent speaker identity. Then we design a VAE-Glow structure, in which the hiding process imperceptibly integrates the source speaker identity into the VC, and the tracing process accurately recovers the source speaker identity and even the source speech in spite of severe speech quality degradation. To address the speech mismatch between the hiding and tracing processes affected by different distortions, we also adopt an asynchronous training strategy to optimize the VAE-Glow models. The VoxTracer is versatile enough to be applied to arbitrary VC methods and popular audio coding standards. Extensive experiments demonstrate that the VoxTracer achieves not only high imperceptibility in hiding, but also nearly 100% tracing accuracy against various types of audio lossy compressions (AAC, MP3, Opus and SILK) with a broad range of bitrates (16 kbps -128 kbps) even in a very short time duration (0.74s). Our speech demo is available at https://anonymous.4open.science/w/DEMOofVoxTracer/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e336349-424f-4dd7-bdd9-fec162bcfa7aCited by top-tier papers4
- Proactive Detection of Voice Cloning with Localized WatermarkingRobin San Roman, Pierre Fernandez, Hady Elsahar, Alexandre Défossez et al.ICML 2024 · 119 citations
- APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic SpeechZhicheng Lian, Lizhi Wang, Hua HuangACM MM 2025 · 1 citation
- De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning AttacksWei Fan, Kejiang Chen, Chang Liu, Weiming Zhang et al.ICML 2025
- AudioMarkNet: Audio Watermarking for Deepfake Speech DetectionWei Zong, Yang-Wai Chow, Willy Susilo, Joonsang Baek et al.USENIX Security 2025
Builds on9
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior et al.ICML 2022 · 602 citations
- HiNet: Deep Image Hiding by Invertible NetworkJunpeng Jing, Xin Deng, Mai Xu, Jianyi Wang et al.ICCV 2021 · 301 citations
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson et al.ICML 2020 · 210 citations
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling SpeakersKaizhi Qian, Yang Zhang, Heting Gao, Junrui Ni et al.ICML 2022 · 157 citations
Related papers
- Catch You and I Can: Revealing Source Voiceprint Against Voice ConversionJiangyi Deng, Yanjiao Chen, Yinan Zhong, Qianhao Miao et al.USENIX Security 2023
- Expose Your Disguise: Recovering Source Speaker Identity From Voice ConversionHanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu et al.CCS 2026
- Tsuikami: Robust Cross-Channel Source Speaker Tracing Under Voice ConversionZhongming Ma, Hanlei Zhang, Shibo Wang, Yanjiao Chen et al.CCS 2026
- ORPHEUS: A Separation-Robust Proactive Defense for Singing Voice ConversionZhaolin Wei, Dengpan Ye, Yanjiao Chen, Jiacheng Deng et al.USENIX Security 2026
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
