Whole-Song Hierarchical Generation of Symbolic Music Using Cascaded Diffusion Models
Ziyu Wang, Lejun Min, Gus Xia
Abstract
Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured whole-song generation. In this paper, we make the first attempt to model a full music piece under the realization of compositional hierarchy. With a focus on symbolic representations of pop songs, we define a hierarchical language, in which each level of hierarchy focuses on the semantics and context dependency at a certain music scope. The high-level languages reveal whole-song form, phrase, and cadence, whereas the low-level languages focus on notes, chords, and their local patterns. A cascaded diffusion model is trained to model the hierarchical language, where each level is conditioned on its upper levels. Experiments and analysis show that our model is capable of generating full-piece music with recognizable global verse-chorus structure and cadences, and the music quality is higher than the baselines. Additionally, we show that the proposed model is controllable in a flexible way. By sampling from the interpretable hierarchical languages or adjusting pre-trained external representations, users can control the music flow via various features such as phrase harmonic structures, rhythmic patterns, and accompaniment texture. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Structured Multi-Track Accompaniment Arrangement via Style Prior ModellingJingwei Zhao, Gus Xia, Ziyu Wang, Ye WangNeurIPS 2024 · 13 citations
- Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured TokenizationLongshen Ou, Jingwei Zhao, Ziyu Wang, Gus Xia et al.NeurIPS 2025 · 5 citations
- Personalized Federated Training of Diffusion Models with Privacy GuaranteesKumar Kshitij Patel, Bingqing Jiang, A. F. M. Mahfuzul Kabir, Weitong Zhang et al.CVPR 2026
- Every Little Bit Helps: Exploring Better Utilization of Unlabeled Data for Semi-supervised Singing Melody Extraction Using Multi-bands Diffusion ModelShuai Yu, Xiaoliang He, Kangjie Dong, Yi YuAAAI 2026
- MuPT: A Generative Symbolic Music Pretrained TransformerXingwei Qu, Yuelin Bai, Yinghao Ma, Ziya Zhou et al.ICLR 2025
Builds on5
- RePaint: Inpainting using Denoising Diffusion Probabilistic ModelsAndreas Lugmayr, Martin Danelljan, Andrés Romero, Fisher Yu et al.CVPR 2022 · 1,425 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- Pop Music Transformer: Beat-based Modeling and Generation of Expressive Pop Piano CompositionsYu-Siang Huang, Yi-Hsuan YangACM MM 2020 · 265 citations
- PopMAG: Pop Music Accompaniment GenerationYi Ren, Jinzheng He, Xu Tan, Tao Qin et al.ACM MM 2020 · 91 citations
- Structured Multi-Track Accompaniment Arrangement via Style Prior ModellingJingwei Zhao, Gus Xia, Ziyu Wang, Ye WangNeurIPS 2024 · 13 citations
Related papers
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 242 citations
- Moûsai: Efficient Text-to-Music Diffusion ModelsFlavio Schneider, Ojasv Kamal, Zhijing Jin, Bernhard SchölkopfACL 2024
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementChenyu Yang, Shuai Wang, Hangting Chen, Wei Tan et al.NeurIPS 2025 · 28 citations
- Structure-Enhanced Pop Music Generation via Harmony-Aware LearningXueyao Zhang, Jinchao Zhang, Yao Qiu, Li Wang et al.ACM MM 2022 · 24 citations
- ExpressiveSinger: Multilingual and Multi-Style Score-based Singing Voice Synthesis with Expressive Performance ControlShuqi Dai, Ming-Yu Liu, Rafael Valle, Siddharth GururaniACM MM 2024 · 6 citations
