Multi-Source Diffusion Models for Simultaneous Music Generation and Separation
Giorgio Mariani, Irene Tallini, Emilian Postolache, Michele Mancusi, Luca Cosmo, Emanuele Rodolà
摘要
In this work, we define a diffusion-based generative model capable of both music synthesis and source separation by learning the score of the joint probability density of sources sharing a context. Alongside the classic total inference tasks (i.e., generating a mixture, separating the sources), we also introduce and experiment on the partial generation task of source imputation, where we generate a subset of the sources given the others (e.g., play a piano track that goes well with the drums). Additionally, we introduce a novel inference method for the separation task based on Dirac likelihood functions. We train our model on Slakh2100, a standard dataset for musical source separation, provide qualitative results in the generation settings, and showcase competitive quantitative results in the source separation setting. Our method is the first example of a single model that can handle both generation and separation tasks, thus representing a step toward general audio models. * Equal contribution. Listing order is random. G.M. wrote most of the code, performed most objective experiments, and contributed to the development of the Dirac separator. I.T. proposed and developed the idea of the Dirac separator and partly formalized its proof, contributed to the code and the objective experiments, especially concerning the Dirac separator, performed the subjective listening tests, and wrote substantial parts of the paper. E.P. proposed and developed the ideas of the source-joint Bayesian separator and the sub-FAD metric, partly formalized the proof of the Dirac separator, contributed to the code and the objective experiments, especially concerning source imputation, and wrote substantial parts of the paper. M.M. proposed the idea of using the source-joint model for music (and accompaniment) generation, proposed using the correction steps, and contributed to the objective experiments. † Shared last authorship.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- Fast Timing-Conditioned Latent Audio DiffusionZach Evans, CJ Carr, Josiah Taylor, Scott H. Hawley 等ICML 2024 · 被引用 220 次
- SAM Audio: Segment Anything in AudioBowen Shi, Andros Tjandra, John Hoffman, Helin Wang 等ICML 2026 · 被引用 35 次
- JEN-1 Composer: A Unified Framework for High-Fidelity Multi-Track Music GenerationYao Yao, Peike Li, Boyu Chen, Alex WangAAAI 2025 · 被引用 19 次
- Score-based Source Separation with Applications to Digital Communication SignalsTejas Jayashankar, Gary C. F. Lee, Alejandro Lancho, Amir Weiss 等NeurIPS 2023 · 被引用 18 次
- Unleashing the Denoising Capability of Diffusion Prior for Solving Inverse ProblemsJiawei Zhang, Jiaxin Zhuang, Cheng Jin, Gen Li 等NeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper14
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao 等ICLR 2021 · 被引用 1,902 次
相关 Paper
- MGE-LDM: Joint Latent Diffusion for Simultaneous Music Generation and Source ExtractionYunkee Chae, Kyogu LeeNeurIPS 2025 · 被引用 4 次
- Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source SeparationShahar Lutati, Eliya Nachmani, Lior WolfICLR 2024 · 被引用 20 次
- A Mixture-Based Framework for Guiding Diffusion ModelsYazid Janati, Badr Moufad, Mehdi Abou El Qassime, Alain Oliviero Durmus 等ICML 2025
- ZeroSep: Separate Anything in Audio with Zero TrainingChao Huang, Yuesheng Ma, Junxuan Huang, Susan Liang 等NeurIPS 2025 · 被引用 8 次
- DITTO: Diffusion Inference-Time T-Optimization for Music GenerationZachary Novack, Julian J. McAuley, Taylor Berg-Kirkpatrick, Nicholas J. BryanICML 2024 · 被引用 81 次
