MusER: Musical Element-Based Regularization for Generating Symbolic Music with Emotion
Shulei Ji, Xinyu Yang
摘要
Generating music with emotion is an important task in automatic music generation, in which emotion is evoked through a variety of musical elements (such as pitch and duration) that change over time and collaborate with each other. However, prior research on deep learning-based emotional music generation has rarely explored the contribution of different musical elements to emotions, let alone the deliberate manipulation of these elements to alter the emotion of music, which is not conducive to fine-grained element-level control over emotions. To address this gap, we present a novel approach employing musical element-based regularization in the latent space to disentangle distinct elements, investigate their roles in distinguishing emotions, and further manipulate elements to alter musical emotions. Specifically, we propose a novel VQ-VAE-based model named MusER. MusER incorporates a regularization loss to enforce the correspondence between the musical element sequences and the specific dimensions of latent variable sequences, providing a new solution for disentangling discrete sequences. Taking advantage of the disentangled latent vectors, a two-level decoding strategy that includes multiple decoders attending to latent vectors with different semantics is devised to better predict the elements. By visualizing latent space, we conclude that MusER yields a disentangled and interpretable latent space and gain insights into the contribution of distinct elements to the emotional dimensions (i.e., arousal and valence). Experimental results demonstrate that MusER outperforms the state-of-the-art models for generating emotional music in both objective and subjective evaluation. Besides, we rearrange music through element transfer and attempt to alter the emotion of music by transferring emotion-distinguishable elements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-LearningDengming Zhang, Weitao You, Ziheng Liu, Lingyun Sun 等AAAI 2025 · 被引用 2 次
- Compose with Me: Collaborative Music Inpainter for Symbolic Music InfillingZhejing Hu, Yan Liu, Gong Chen, Bruce X. B. YuAAAI 2025 · 被引用 2 次
- Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsZhejing Hu, Yan Liu, Zhi Zhang, Aiwei Zhang 等AAAI 2026
它引用的顶会 Paper4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 被引用 242 次
- PopMAG: Pop Music Accompaniment GenerationYi Ren, Jinzheng He, Xu Tan, Tao Qin 等ACM MM 2020 · 被引用 91 次
相关 Paper
- Emotion-Based End-to-End Matching Between Image and Music in Valence-Arousal SpaceSicheng Zhao, Yaxian Li, Xingxu Yao, Weizhi Nie 等ACM MM 2020 · 被引用 30 次
- Music2Palette: Emotion-aligned Color Palette Generation via Cross-Modal Representation LearningJiayun Hu, Yueyi He, Tianyi Liang, Changbo Wang 等ACM MM 2025 · 被引用 2 次
- Sera: Separated Coarse-to-fine Representation Alignment for Cross-subject EEG-based Emotion RecognitionZhihao Jia, Meiyan Xu, Jingyuan Wang, Ziyu Jia 等ACM MM 2025 · 被引用 2 次
- Is Discourse Role Important for Emotion Recognition in Conversation?Donovan Ong, Jian Su, Bin Chen, Anh Tuan Luu 等AAAI 2022 · 被引用 29 次
- Emotional Voice PuppetryYe Pan, Ruisi Zhang, Shengran Cheng, Shuai Tan 等IEEE VR 2023 · 被引用 21 次
