Phoneme Hallucinator: One-Shot Voice Conversion via Set Expansion
Siyuan Shan, Yang Li, Amartya Banerjee, Junier B. Oliva
Abstract
Voice conversion (VC) aims at altering a person's voice to make it sound similar to the voice of another person while preserving linguistic content. Existing methods suffer from a dilemma between content intelligibility and speaker similarity; i.e., methods with higher intelligibility usually have a lower speaker similarity, while methods with higher speaker similarity usually require plenty of target speaker voice data to achieve high intelligibility. In this work, we propose a novel method Phoneme Hallucinator that achieves the best of both worlds. Phoneme Hallucinator is a one-shot VC model; it adopts a novel model to hallucinate diversified and high-fidelity target speaker phonemes based just on a short target speaker voice (e.g. 3 seconds). The hallucinated phonemes are then exploited to perform neighbor-based voice conversion. Our model is a text-free, any-to-any VC model that requires no text annotations and supports conversion to any unseen speaker. Quantitative and qualitative evaluations show that Phoneme Hallucinator outperforms existing VC methods for both intelligibility and speaker similarity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30dc2d0f-6503-4f76-871d-7f86243eb70eCited by top-tier papers1
Ask how each one uses itBuilds on8
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 2,890 citations
- YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for EveryoneEdresson Casanova, Julian Weber, Christopher Dane Shulby, Arnaldo Cândido Júnior et al.ICML 2022 · 602 citations
- Empower Entity Set Expansion via Language Model ProbingYunyi Zhang, Jiaming Shen, Jingbo Shang, Jiawei HanACL 2020 · 51 citations
- Exchangeable Neural ODE for Set ModelingYang Li, Haidong Yi, Christopher M. Bender, Siyuan Shan et al.NeurIPS 2020 · 32 citations
Related papers
- PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice ConversionYimin Deng, Huaizhen Tang, Xulong Zhang, Jianzong Wang et al.ACM MM 2023 · 15 citations
- Improving Zero-Shot Voice Style Transfer via Disentangled Representation LearningSiyang Yuan, Pengyu Cheng, Ruiyi Zhang, Weituo Hao et al.ICLR 2021 · 64 citations
- Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling SchemeVadim Popov, Ivan Vovk, Vladimir Gogoryan, Tasnima Sadekova et al.ICLR 2022 · 185 citations
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 5 citations
- StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow MatchingJixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning et al.AAAI 2025 · 13 citations
