Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion
Hanlei Zhang, Zhongming Ma, Mingyang Zhang, Tengfei Liu, Yushi Cheng, Yanjiao Chen
摘要
Voice conversion (VC) poses a significant threat to biometric security by allowing attackers to impersonate target speakers. In forensic contexts, recovering the source speaker's identity from converted audio is vital for narrowing the field of suspects. To address this, we propose TRIDENT, a retracing framework designed to restore a source speaker's original identity from a converted audio sample. TRIDENT utilizes a three-pronged architecture consisting of a primary extractor and two auxiliary branches. The first auxiliary branch identifies the underlying voice conversion mechanism. This design acknowledges that even if the exact conversion strategy is unknown, a high-performance model adopted by the attacker is typically a derivative or variant of established mainstream ones. The second auxiliary branch extracts a latent representation of the target speaker, facilitating the isolation of target-specific traits from the composite converted audio sample. Finally, the main extractor leverages insights from both auxiliary branches to decouple confounding factors and distill a highly discriminative representation of the source speaker's identity. Experimental results demonstrate that TRIDENT achieves an accuracy as high as 90.99% against 7 state-of-the-art voice conversion methods. Furthermore, TRIDENT maintains robust performance under challenging conditions, including telephony channels, unseen languages, and adaptive scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling SchemeVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova 等ICLR 2022 · 被引用 185 次
- DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice ConversionHa-Yeong Choi, Sang-Hoon Lee, Seong-Whan LeeAAAI 2024 · 被引用 66 次
- VoiceRadar: Voice Deepfake Detection using Micro-Frequency and Compositional AnalysisKavita Kumari, Maryam Abbasihafshejani, Alessandro Pegoraro, Phillip Rieger 等NDSS 2025
- Catch You and I Can: Revealing Source Voiceprint Against Voice ConversionJiangyi Deng, Yanjiao Chen, Yinan Zhong, Qianhao Miao 等USENIX Security 2023
相关 Paper
- Who is Speaking Actually? Robust and Versatile Speaker Traceability for Voice ConversionYanzhen Ren, Hongcheng Zhu, Liming Zhai, Zongkun Sun 等ACM MM 2023 · 被引用 3 次
- Tsuikami: Robust Cross-Channel Source Speaker Tracing Under Voice ConversionZhongming Ma, Hanlei Zhang, Shibo Wang, Yanjiao Chen 等CCS 2026
- ORPHEUS: A Separation-Robust Proactive Defense for Singing Voice ConversionZhaolin Wei, Dengpan Ye, Yanjiao Chen, Jiacheng Deng 等USENIX Security 2026
- Utilizing Speaker Profiles for Impersonation Audio DetectionHao Gu, Jiangyan Yi, Chenglong Wang, Yong Ren 等ACM MM 2024 · 被引用 3 次
- Parrot-Trained Adversarial Examples: Pushing the Practicality of Black-Box Audio Attacks against Speaker Recognition ModelsRui Duan, Zhe Qu, Leah Ding, Yao Liu 等NDSS 2024
