EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
Ashishkumar Prabhakar Gudmalwar, Ishan Darshan Biyani, Nirmesh J. Shah, Pankaj Wasnik, Rajiv Ratn Shah
摘要
The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion intensity in the diffusion-based EVC framework to generate precise speech of the target emotion. Traditional approaches control the intensity of an emotional state in the utterance via emotion class probabilities or intensity labels that often lead to inept style manipulations and degradations in quality. On the contrary, we aim to regulate emotion intensity using selfsupervised learning-based feature representations and unsupervised directional latent vector modeling (DVM) in the emotional embedding space within a diffusion-based framework. These emotion embeddings can be modified based on the given target emotion intensity and the corresponding direction vector. Furthermore, the updated embeddings can be fused in the reverse diffusion process to generate the speech with the desired emotion and intensity. In summary, this paper aims to achieve high-quality emotional intensity regularization in the diffusion-based EVC framework, which is the first of its kind work. The effectiveness of the proposed method has been shown across state-of-the-art (SOTA) baselines in terms of subjective and objective evaluations for the English and Hindi languages 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman 等ICML 2023 · 被引用 6,966 次
- Neural Dubber: Dubbing for Videos According to ScriptsChenxu Hu, Qiao Tian, Tingle Li, Yuping Wang 等NeurIPS 2021 · 被引用 62 次
- VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingYihan Wu, Junliang Guo, Xu Tan, Chen Zhang 等AAAI 2023 · 被引用 33 次
相关 Paper
- DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice ConversionHa-Yeong Choi, Sang-Hoon Lee, Seong-Whan LeeAAAI 2024 · 被引用 66 次
- Audio-Driven Emotional Video PortraitsXinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu 等CVPR 2021
- Textless Speech Emotion Conversion using Discrete & Decomposed RepresentationsFelix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov 等EMNLP 2022 · 被引用 25 次
- Emotional Face-to-SpeechJiaxin Ye, Boyuan Cao, Hongming ShanICML 2025
- Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-SpeechFangming Feng, Dongjie Fu, Zequn Xie, Yu Zhang 等ACL 2026
