JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music Generation
Yao Yao, Peike Li, Boyu Chen, Alex Wang
Abstract
With rapid advances in generative artificial intelligence, the text-to-music synthesis task has emerged as a promising direction for music generation. Nevertheless, achieving precise control over multi-track generation remains an open challenge. While existing models excel in directly generating multi-track mix, their limitations become evident when it comes to composing individual tracks and integrating them in a controllable manner. This departure from the typical workflows of professional composers hinders the ability to refine details in specific tracks. To address this gap, we propose JEN-1 Composer, a unified framework designed to efficiently model marginal, conditional, and joint distributions over multi-track music using a single model. Building upon an audio latent diffusion model, JEN-1 Composer extends the versatility of multi-track music generation. We introduce a progressive curriculum training strategy, which gradually escalates the difficulty of training tasks while ensuring the model's generalization ability and facilitating smooth transitions between different scenarios. During inference, users can iteratively generate and select music tracks, thus incrementally composing entire musical pieces in accordance with the Human-AI co-composition workflow. Our approach demonstrates state-of-the-art performance in controllable and highfidelity multi-track music synthesis, marking a significant advancement in interactive AI-assisted music creation. Our demo pages are available at www.jenmusic.ai/research. Introduction The rapid evolution of generative modeling has positioned AI-driven music generation as a prominent field, merging research innovation with practical applications in the music industry. Early systems like Music Transformer (Huang et al. 2018) and MuseNet (Payne 2019), which utilized symbolic representations (Engel et al. 2017) , were pivotal in translating textual descriptions into MIDI-style outputs. Although these methods were groundbreaking, their dependence on predefined virtual synthesizers often compromised the audio quality and restricted the diversity of their musical outputs. Recent advancements in text-to-music synthesis, as demonstrated by models like MusicGen (Copet et al. 2024), MusicLM (Agostinelli et al. 2023), and Jen-1 (Li et al.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc0fc430-8b60-4794-a551-3420066cd9c0Cited by top-tier papers6
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source ExtractionYunkee Chae, Kyogu LeeNeurIPS 2025 · 4 citations
- JEN-1 DreamStyler: Customized Musical Concept Learning via Pivotal Parameters TuningBoyu Chen, Peike Li, Yao Yao, Alex WangAAAI 2025 · 3 citations
- SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music GenerationHongrui Wang, Fan Zhang, Zhiyuan Yu, Ziya Zhou et al.ICLR 2026
- Is Symbolic Music a Specific Language? Exploring Inspiration-to-Structure Machine Composition via LLMsZhejing Hu, Yan Liu, Zhi Zhang, Aiwei Zhang et al.AAAI 2026
Builds on12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- Simple and Controllable Music GenerationJade Copet, Felix Kreuk, Itai Gat, Tal Remez et al.NeurIPS 2023 · 843 citations
- AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsHaohe Liu, Zehua Chen, Yi Yuan, Xinhao Mei et al.ICML 2023 · 773 citations
- Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion ModelsRongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren et al.ICML 2023 · 469 citations
Related papers
- MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music CompositionPhilippe Pasquier, Jeff Ens, Nathan Fradet, Paul Triana et al.AAAI 2025 · 14 citations
- Multimodal Large Language Models for Multi-Subject In-Context Image GenerationYucheng Zhou, Dubing Chen, Huan Zheng, Jianbing ShenACL 2026 · 2 citations
- UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsGuozhen Zhang, Zixiang Zhou, Teng Hu, Ziqiao Peng et al.CVPR 2026 · 40 citations
- SongComposer: A Large Language Model for Lyric and Melody Generation in Song CompositionShuangrui Ding, Zihan Liu, Xiaoyi Dong, Pan Zhang et al.ACL 2025
- Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and EditingZeyue Tian, Binxin Yang, Zhaoyang Liu, Jiexuan Zhang et al.SIGGRAPH 2026
