Composition of Pretrained Diffusion Models: A Logic-Based Calculus
Peter Blohm, Vikas K Garg
摘要
Composing pretrained diffusion models provides a cost-effective mechanism for encoding constraints and unlocking complex generative capabilities.
Prior work relies on crafting compositional operators that seek to extend set-theoretic notions such as union and intersection to diffusion models, e.g., using a product or mixture of the underlying energy functions.
We expose the inadequacy and inconsistency of combining these operators, including limited mode coverage, biased sampling, instability under negation queries, and failure to satisfy basic compositional laws such as idempotency and distributivity.
We introduce a principled calculus grounded in fuzzy logic that resolves these issues.
Specifically, we define a general class of conjunction, disjunction, and negation operators that generalize the classical mixture-, product-, and harmonic-mean-style operators, illustrating how they circumvent various pathologies and enable precise combinatorial reasoning with score models.
Beyond existing methods, the proposed Dombi operators yield complex generative outcomes, such as XOR-style logical compositions of pretrained score models.
We establish rigorous theoretical guarantees on the stability of Dombi compositions, and derive Feynman-Kac correctors to mitigate the sampling bias in score composition.
Empirical results on image generation with Stable Diffusion and multi-objective molecular generation substantiate the conceptual, theoretical, and methodological benefits.
Overall, this work lays the foundation for systematic design, analysis, and deployment of diffusion ensembles.
Code is available at github.com/Aalto-QuML/logic-diffusion-composition.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
相关 Paper
- Logical Guidance for the Exact Composition of Diffusion ModelsFrancesco Alesiani, Jonathan Warrell, Tanja Bien, Henrik Christiansen 等ICML 2026
- The Superposition of Diffusion Models Using the Itô Density EstimatorMarta Skreta, Lazar Atanackovic, Joey Bose, Alexander Tong 等ICLR 2025
- Mechanisms of Projective Composition of Diffusion ModelsArwen Bradley, Preetum Nakkiran, David Berthelot, James Thornton 等ICML 2025
- Reduce, Reuse, Recycle: Compositional Generation with Energy-Based Diffusion Models and MCMCYilun Du, Conor Durkan, Robin Strudel, Joshua B. Tenenbaum 等ICML 2023 · 被引用 219 次
- Score Correction for Generative Models with Probabilistic ConstraintsShishang Wu, Bingjing Tang, Vinayak A RaoICML 2026
