StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
Zhichao Wang, Yuanzhe Chen, Xinsheng Wang, Lei Xie, Yuping Wang
摘要
Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance. However, existing LM-based VC models usually apply offline conversion from source semantics to acoustic features, demanding the complete source speech and limiting their deployment to realtime applications. In this paper, we introduce StreamVoice, a novel streaming LM-based model for zero-shot VC, facilitating real-time conversion given arbitrary speaker prompts and source speech. Specifically, to enable streaming capability, StreamVoice employs a fully causal context-aware LM with a temporalindependent acoustic predictor, while alternately processing semantic and acoustic features at each time step of autoregression which eliminates the dependence on complete source speech. To address the potential performance degradation from the incomplete context in streaming processing, we enhance the contextawareness of the LM through two strategies: 1) teacher-guided context foresight, using a teacher model to summarize the present and future semantic context during training to guide the model's forecasting for missing context; 2) semantic masking strategy, promoting acoustic prediction from preceding corrupted semantic and acoustic input, enhancing context-learning ability. Notably, StreamVoice is the first LMbased streaming zero-shot VC model without any future look-ahead. Experiments demonstrate StreamVoice's streaming conversion capability while achieving zero-shot performance comparable to non-streaming VC systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper4
- NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing SynthesizersKai Shen, Zeqian Ju, Xu Tan, Eric Liu 等ICLR 2024 · 被引用 362 次
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised RepresentationsHyeong-Seok Choi, Juheon Lee, Wansoo Kim, Jie Lee 等NeurIPS 2021 · 被引用 200 次
- Retriever: Learning Content-Style Representation as a Token-Level Bipartite GraphDacheng Yin, Xuanchi Ren, Chong Luo, Yuwang Wang 等ICLR 2022 · 被引用 13 次
- Revisiting Over-Smoothness in Text to SpeechYi Ren, Xu Tan, Tao Qin, Zhou Zhao 等ACL 2022
相关 Paper
- StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow MatchingJixun Yao, Yuguang Yang, Yu Pan, Ziqian Ning 等AAAI 2025 · 被引用 13 次
- SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech TranslationKeqi Deng, Wenxi Chen, Xie Chen, Philip C. WoodlandACL 2025
- Face-Driven Zero-Shot Voice Conversion with Memory-based Face-Voice AlignmentZhengyan Sheng, Yang Ai, Yan-Nian Chen, Zhen-Hua LingACM MM 2023 · 被引用 5 次
- Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre ModelingYuguang Yang, Yu Pan, Jixun Yao, Xiang Zhang 等ACL 2025
- Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural TransducersGrant P. Strimel, Yi Xie, Brian John King, Martin Radfar 等ICML 2023 · 被引用 12 次
