Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
Ziyu Wang, Lejun Min, Gus Xia
摘要
Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured whole-song generation. In this paper, we make the first attempt to model a full music piece under the realization of compositional hierarchy. With a focus on symbolic representations of pop songs, we define a hierarchical language, in which each level of hierarchy focuses on the semantics and context dependency at a certain music scope. The high-level languages reveal whole-song form, phrase, and cadence, whereas the low-level languages focus on notes, chords, and their local patterns. A cascaded diffusion model is trained to model the hierarchical language, where each level is conditioned on its upper levels. Experiments and analysis show that our model is capable of generating full-piece music with recognizable global verse-chorus structure and cadences, and the music quality is higher than the baselines. Additionally, we show that the proposed model is controllable in a flexible way. By sampling from the interpretable hierarchical languages or adjusting pre-trained external representations, users can control the music flow via various features such as phrase harmonic structures, rhythmic patterns, and accompaniment texture. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Structured Multi-Track Accompaniment Arrangement via Style Prior ModellingJingwei Zhao, Gus Xia, Ziyu Wang, Ye WangNeurIPS 2024 · 被引用 13 次
- Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured TokenizationLongshen Ou, Jingwei Zhao, Ziyu Wang, Gus Xia 等NeurIPS 2025 · 被引用 5 次
- Personalized Federated Training of Diffusion Models with Privacy GuaranteesKumar Kshitij Patel, Bingqing Jiang, A. F. M. Mahfuzul Kabir, Weitong Zhang 等CVPR 2026
- Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion ModelShuai Yu, Xiaoliang He, Kangjie Dong, Yi YuAAAI 2026
- MuPT: A Generative Symbolic Music Pretrained TransformerXingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou 等ICLR 2025
它引用的顶会 Paper5
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu 等CVPR 2022 · 被引用 1,425 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 被引用 265 次
- PopMAG: Pop Music Accompaniment GenerationYi Ren, Jinzheng He, Xu Tan, Tao Qin 等ACM MM 2020 · 被引用 91 次
- Structured Multi-Track Accompaniment Arrangement via Style Prior ModellingJingwei Zhao, Gus Xia, Ziyu Wang, Ye WangNeurIPS 2024 · 被引用 13 次
相关 Paper
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 被引用 242 次
- Moûsai: Efficient Text-to-Music Diffusion ModelsFlavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard SchölkopfACL 2024
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementChenyu Yang, Shuai Wang, Hangting Chen, Wei Tan 等NeurIPS 2025 · 被引用 28 次
- Structure-Enhanced Pop Music Generation via Harmony-Aware LearningXueyao Zhang, Jinchao Zhang, Yao Qiu, Li Wang 等ACM MM 2022 · 被引用 24 次
- ExpressiveSinger: Multilingual and Multi-Style Score-based Singing Voice Synthesis with Expressive Performance ControlShuqi Dai, Ming-Yu Liu, Rafael Valle, Siddharth GururaniACM MM 2024 · 被引用 6 次
