YuE: Scaling Open Foundation Models for Long-Form Music Generation
Ruibin Yuan, Hanfeng Lin, Shuyue Guo, Ge Zhang, Jiahao Pan, Yongyi Zang, Haohe Liu, Yiming Liang, Wenye Ma, Xingjian Du, Xeron Du, Zhen Ye
摘要
We tackle the task of long-form music generation-particularly the challenging lyrics-to-song problem-by introducing YuE (乐), a family of open foundation models based on the LLaMA2 architecture. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while maintaining lyrical alignment, coherent musical structure, and engaging vocal melodies with appropriate accompaniment. It achieves this through: (1) track-decoupled nexttoken prediction to overcome dense mixture signals, (2) structural progressive conditioning for long-context lyrical alignment, and (3) a multitask, multiphase pre-training recipe to converge and generalize. In addition, we redesign the in-context learning technique for music generation, enabling versatile style transfer (e.g., converting Japanese city pop into an English rap while preserving the original accompaniment) and bidirectional generation. Through extensive evaluation, we demonstrate that YuE matches or even surpasses some of the proprietary systems in musicality and vocal agility. In addition, fine-tuning YuE enables additional controls and enhanced support for tail languages. Furthermore, beyond generation, we show that YuE's learned representations can perform competatively on music understanding tasks, where the results of YuE match or exceed state-of-the-art methods on the MARBLE benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- LeVo: High-Quality Song Generation with Multi-Preference AlignmentShun Lei, Yaoxun Xu, Zhiwei Lin, Huaicheng Zhang 等NeurIPS 2025 · 被引用 43 次
- SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementChenyu Yang, Shuai Wang, Hangting Chen, Wei Tan 等NeurIPS 2025 · 被引用 28 次
- Audio Super-Resolution with Latent Bridge ModelsChang Li, Zehua Chen, Liyuan Wang, Jun ZhuNeurIPS 2025 · 被引用 18 次
- UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text InstructionsChunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang 等ACL 2026 · 被引用 2 次
- Harmonic Canvas: Inversion-Free Editing for Visually-Guided Music Style TransferYue Lei, Siqi Yang, Ting Zhong, Fan ZhouCVPR 2026
它引用的顶会 Paper15
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 被引用 9,451 次
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and LanguageAlexei Baevski, Wei-Ning Hsu, Qiantong Xu, Arun Babu 等ICML 2022 · 被引用 1,123 次
- High-Fidelity Audio Compression with Improved RVQGANRithesh Kumar, Prem Seetharaman, Alejandro Luebs, Ishaan Kumar 等NeurIPS 2023 · 被引用 910 次
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez 等NeurIPS 2023 · 被引用 843 次
相关 Paper
- SongGLM: Lyric-to-Melody Generation with 2D Alignment Encoding and Multi-Task Pre-TrainingJiaxing Yu, Xinda Wu, Yunfei Xu, Tieyao Zhang 等AAAI 2025 · 被引用 2 次
- CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical ControlsLi Chai, Donglin WangAAAI 2025 · 被引用 1 次
- SegTune: Structured and Fine-Grained Control for Song GenerationYuejiao Wang, Zihao Ji, Pengfei Cai, Xu Li 等ACL 2026 · 被引用 2 次
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed HypergraphsWen-Yi Hsiao, Jen-Yu Liu, Yin-Cheng Yeh, Yi-Hsuan YangAAAI 2021 · 被引用 242 次
- SongMASS: Automatic Song Writing with Pre-training and Alignment ConstraintZhonghao Sheng, Kaitao Song, Xu Tan, Yi Ren 等AAAI 2021 · 被引用 84 次
