Seeing Your Speech Style: A Novel Zero-Shot Identity-Disentanglement Face-based Voice Conversion
Yan Rong, Li Liu
Abstract
Face-based Voice Conversion (FVC) is a novel task that leverages facial images to generate the target speaker's voice style. Previous work has two shortcomings: (1) suffering from obtaining facial embeddings that are well-aligned with the speaker's voice identity information, and (2) inadequacy in decoupling content and speaker identity information from the audio input. To address these issues, we present a novel FVC method, Identity-Disentanglement Face-based Voice Conversion (ID-FaceVC), which overcomes the above two limitations. More precisely, we propose an Identity-Aware Query-based Contrastive Learning (IAQ-CL) module to extract speaker-specific facial features, and a Mutual Information-based Dual Decoupling (MIDD) module to purify content features from audio, ensuring clear and high-quality voice conversion. Besides, unlike prior works, our method can accept either audio or text inputs, offering controllable speech generation with adjustable emotional tone and speed. Extensive experiments demonstrate that ID-FaceVC achieves state-ofthe-art performance across various metrics, with qualitative and user study results confirming its effectiveness in naturalness, similarity, and diversity. Project website with audio samples and code can be found at https://id-facevc.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64a7c6f8-dc2c-47b8-80a2-fa04f8f9bd4fCited by top-tier papers4
- Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyTianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang et al.EMNLP 2025 · 10 citations
- AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio GenerationYan Rong, Jinting Wang, Guangzhi Lei, Shan Yang et al.ACM MM 2025 · 1 citation
- Visual-informed Silent Video Identity ConversionYifan Liu, Yu Fang, Zhouhan LinACM MM 2025
- Coordinated Disentanglement with Iterative Mode Discovery Under Hidden CorrelationsRong Hu, Ling ChenICML 2026
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-SpeechJaehyeon Kim, Jungil Kong, Juhee SonICML 2021 · 1,267 citations
- CLUB: A Contrastive Log-ratio Upper Bound of Mutual InformationPengyu Cheng, Weituo Hao, Shuyang Dai, Jiachang Liu et al.ICML 2020 · 512 citations
Related papers
- Zero-Shot Face-Based Voice Conversion: Bottleneck-Free Speech Disentanglement in the Real-World ScenarioShao-En Weng, Hong-Han Shuai, Wen-Huang ChengAAAI 2023 · 4 citations
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- Face-based Voice Conversion: Learning the Voice behind a FaceHsiao-Han Lu, Shao-En Weng, Ya-Fan Yen, Hong-Han Shuai et al.ACM MM 2021 · 15 citations
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 5 citations
- StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow MatchingJixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning et al.AAAI 2025 · 13 citations
