TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna
摘要
Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varying, while speaker identity is injected as a static global embedding. We introduce a streamable speech synthesizer that aligns the temporal granularity of identity and content via a content-synchronous, time-varying timbre (TVT) representation. A Global Timbre Memory expands a global timbre instance into multiple compact facets; frame-level content attends to this memory, a gate regulates variation, and spherical interpolation preserves identity geometry while enabling smooth local changes. In addition, a factorized vector-quantized bottleneck regularizes content to reduce residual speaker leakage. The resulting system is streamable end-to-end, with <80 ms GPU latency. Experiments show improvements in naturalness, speaker transfer, and anonymization compared to SOTA streaming baselines, establishing TVT as a scalable approach for privacy-preserving and expressive speech synthesis under strict latency budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion ModelsZeqian Ju, Yuancheng Wang, Kai Shen, Xu Tan 等ICML 2024 · 被引用 341 次
- SlerpFace: Face Template Protection via Spherical Linear InterpolationZhizhou Zhong, Yuxi Mi, Yuge Huang, Jianqing Xu 等AAAI 2025 · 被引用 14 次
相关 Paper
- CSSinger: End-to-End Chunkwise Streaming Singing Voice Synthesis System Based on Conditional Variational AutoencoderJianwei Cui, Yu Gu, Shihao Chen, Jie Zhang 等AAAI 2025 · 被引用 1 次
- Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow MatchingJialong Zuo, Shengpeng Ji, Minghui Fang, Mingze Li 等ACL 2025 · 被引用 1 次
- V-Cloak: Intelligibility-, Naturalness- & Timbre-Preserving Real-Time Voice AnonymizationJiangyi Deng, Fei Teng, Yanjiao Chen, Xiaofu Chen 等USENIX Security 2023
- Unsupervised Audiovisual Synthesis via Exemplar AutoencodersKangle Deng, Aayush Bansal, Deva RamananICLR 2021 · 被引用 4 次
- VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio ModelsPatrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison 等NeurIPS 2022 · 被引用 18 次
