EmoReg: Directional Latent Vector Modeling for Emotional Intensity Regularization in Diffusion-based Voice Conversion
Ashishkumar Prabhakar Gudmalwar, Ishan Darshan Biyani, Nirmesh J. Shah, Pankaj Wasnik, Rajiv Ratn Shah
Abstract
The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion intensity in the diffusion-based EVC framework to generate precise speech of the target emotion. Traditional approaches control the intensity of an emotional state in the utterance via emotion class probabilities or intensity labels that often lead to inept style manipulations and degradations in quality. On the contrary, we aim to regulate emotion intensity using selfsupervised learning-based feature representations and unsupervised directional latent vector modeling (DVM) in the emotional embedding space within a diffusion-based framework. These emotion embeddings can be modified based on the given target emotion intensity and the corresponding direction vector. Furthermore, the updated embeddings can be fused in the reverse diffusion process to generate the speech with the desired emotion and intensity. In summary, this paper aims to achieve high-quality emotional intensity regularization in the diffusion-based EVC framework, which is the first of its kind work. The effectiveness of the proposed method has been shown across state-of-the-art (SOTA) baselines in terms of subjective and objective evaluations for the English and Hindi languages 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b74373f-3dac-4751-9b9c-357b037a680aBuilds on3
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- Neural Dubber: Dubbing for Videos According to ScriptsChenxu Hu, Qiao Tian, Tingle Li, Yuping Wang et al.NeurIPS 2021 · 62 citations
- VideoDubber: Machine Translation with Speech-Aware Length Control for Video DubbingYihan Wu, Junliang Guo, Xu Tan, Chen Zhang et al.AAAI 2023 · 33 citations
Related papers
- DDDM-VC: Decoupled Denoising Diffusion Models with Disentangled Representation and Prior Mixup for Verified Robust Voice ConversionHa-Yeong Choi, Sang-Hoon Lee, Seong-Whan LeeAAAI 2024 · 66 citations
- Audio-Driven Emotional Video PortraitsXinya Ji, Hang Zhou, Kaisiyuan Wang, Wayne Wu et al.CVPR 2021
- Textless Speech Emotion Conversion using Discrete & Decomposed RepresentationsFelix Kreuk, Adam Polyak, Jade Copet, Eugene Kharitonov et al.EMNLP 2022 · 25 citations
- Emotional Face-to-SpeechJiaxin Ye, Boyuan Cao, Hongming ShanICML 2025
- Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-SpeechFangming Feng, Dongjie Fu, Zequn Xie, Yu Zhang et al.ACL 2026
