Textless Speech Emotion Conversion using Discrete & Decomposed Representations
Felix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov, Tu Anh Nguyen, Morgane Rivière, Wei-Ning Hsu, Abdelrahman Mohamed, Emmanuel Dupoux, Yossi Adi
摘要
Speech emotion conversion is the task of modifying the perceived emotion of a speech utterance while preserving the lexical content and speaker identity. In this study, we cast the problem of emotion conversion as a spoken language translation task. We use a decomposition of the speech signal into discrete learned representations, consisting of phonetic-content units, prosodic features, speaker, and emotion. First, we modify the speech content by translating the phoneticcontent units to a target emotion, and then predict the prosodic features based on these units. Finally, the speech waveform is generated by feeding the predicted representations into a neural vocoder. Such a paradigm allows us to go beyond spectral and parametric changes of the signal, and model non-verbal vocalizations, such as laughter insertion, yawning removal, etc. We demonstrate objectively and subjectively that the proposed method is vastly superior to current approaches and even beats text-based systems in terms of perceived emotion and audio quality. We rigorously evaluate all components of such a complex system and conclude with an extensive model analysis and ablation study to better emphasize the architectural choices, strengths and weaknesses of the proposed method. Samples are available under the following link: [samples].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- AudioGen: Textually Guided Audio GenerationFelix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer 等ICLR 2023 · 被引用 54 次
- From Discrete Tokens to High-Fidelity Audio Using Multi-Band DiffusionRobin San Roman, Yossi Adi, Antoine Deleforge, Romain Serizel 等NeurIPS 2023 · 被引用 50 次
- SelfVC: Voice Conversion With Iterative Refinement using Self TransformationsPaarth Neekhara, Shehzeen Samarah Hussain, Rafael Valle, Boris Ginsburg 等ICML 2024 · 被引用 7 次
它引用的顶会 Paper5
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech SynthesisJungil Kong, Jaehyeon Kim, Jaekyoung BaeNeurIPS 2020 · 被引用 2,890 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- FastSpeech 2: Fast and High-Quality End-to-End Text to SpeechYi Ren, Chenxu Hu, Xu Tan, Tao Qin 等ICLR 2021 · 被引用 513 次
- Unsupervised Speech Decomposition via Triple Information BottleneckKaizhi Qian, Yang Zhang, Shiyu Chang, Mark Hasegawa-Johnson 等ICML 2020 · 被引用 210 次
相关 Paper
- Neural Emotion Director: Speech-preserving semantic control of facial expressions in "in-the-wild" videosFoivos Paraperas Papantoniou, Panagiotis Paraskevas Filntisis, Petros Maragos, Anastasios RoussosCVPR 2022 · 被引用 32 次
- Cross-Modal Emotion Transfer for Emotion Editing in Talking Face VideoChanhyuk Choi, Taesoo Kim, Donggyu Lee, Siyeol Jung 等CVPR 2026 · 被引用 1 次
- PMVC: Data Augmentation-Based Prosody Modeling for Expressive Voice ConversionYimin Deng, Huaizhen Tang, Xulong Zhang, Jianzong Wang 等ACM MM 2023 · 被引用 15 次
- EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationZiqiao Peng, Haoyu Wu, Zhenbo Song, Hao Xu 等ICCV 2023 · 被引用 192 次
- SECap: Speech Emotion Captioning with Large Language ModelYaoxun Xu, Hangting Chen, Jianwei Yu, Qiaochu Huang 等AAAI 2024 · 被引用 70 次
