Multi-Source Diffusion Models for Simultaneous Music Generation and Separation
Giorgio Mariani, Irene Tallini, Emilian Postolache, Michele Mancusi, Luca Cosmo, Emanuele Rodolà
Abstract
In this work, we define a diffusion-based generative model capable of both music synthesis and source separation by learning the score of the joint probability density of sources sharing a context. Alongside the classic total inference tasks (i.e., generating a mixture, separating the sources), we also introduce and experiment on the partial generation task of source imputation, where we generate a subset of the sources given the others (e.g., play a piano track that goes well with the drums). Additionally, we introduce a novel inference method for the separation task based on Dirac likelihood functions. We train our model on Slakh2100, a standard dataset for musical source separation, provide qualitative results in the generation settings, and showcase competitive quantitative results in the source separation setting. Our method is the first example of a single model that can handle both generation and separation tasks, thus representing a step toward general audio models. * Equal contribution. Listing order is random. G.M. wrote most of the code, performed most objective experiments, and contributed to the development of the Dirac separator. I.T. proposed and developed the idea of the Dirac separator and partly formalized its proof, contributed to the code and the objective experiments, especially concerning the Dirac separator, performed the subjective listening tests, and wrote substantial parts of the paper. E.P. proposed and developed the ideas of the source-joint Bayesian separator and the sub-FAD metric, partly formalized the proof of the Dirac separator, contributed to the code and the objective experiments, especially concerning source imputation, and wrote substantial parts of the paper. M.M. proposed the idea of using the source-joint model for music (and accompaniment) generation, proposed using the correction steps, and contributed to the objective experiments. † Shared last authorship.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a74e545e-3145-4c7f-abf8-50ec886ff907Cited by top-tier papers16
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley et al.ICML 2024 · 220 citations
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang et al.ICML 2026 · 35 citations
- JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music GenerationYao Yao, Peike Li, Boyu Chen, Alex WangAAAI 2025 · 19 citations
- Score-based Source Separation with Applications to Digital Communication SignalsTejas Jayashankar, Gary C. F. Lee, Alejandro Lancho, Amir Weiss et al.NeurIPS 2023 · 18 citations
- Unleashing the Denoising Capability of Diffusion Prior for Solving Inverse ProblemsJiawei Zhang, Jiaxin Zhuang, Cheng Jin, Gen Li et al.NeurIPS 2024 · 11 citations
Builds on14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
Related papers
- MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source ExtractionYunkee Chae, Kyogu LeeNeurIPS 2025 · 4 citations
- Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source SeparationShahar Lutati, Eliya Nachmani, Lior WolfICLR 2024 · 20 citations
- A Mixture-Based Framework for Guiding Diffusion ModelsYazid Janati, Badr Moufad, Mehdi Abou El Qassime, Alain Oliviero Durmus et al.ICML 2025
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang et al.NeurIPS 2025 · 8 citations
- DITTO: Diffusion Inference-Time T-Optimization for Music GenerationZachary Novack, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. BryanICML 2024 · 81 citations
