On improved Conditioning Mechanisms and Pre-training Strategies for Diffusion Models
Tariq Berrada Ifriqi, Pietro Astolfi, Melissa Hall, Reyhane Askari Hemmat, Yohann Benchetrit, Marton Havasi, Matthew J. Muckley, Karteek Alahari, Adriana Romero-Soriano, Jakob Verbeek, Michal Drozdzal
Abstract
Large-scale training of latent diffusion models (LDMs) has enabled unprecedented quality in image generation. However, the key components of the best performing LDM training recipes are oftentimes not available to the research community, preventing apple-to-apple comparisons and hindering the validation of progress in the field. In this work, we perform an in-depth study of LDM training recipes focusing on the performance of models and their training efficiency. To ensure apple-to-apple comparisons, we re-implement five previously published models with their corresponding recipes. Through our study, we explore the effects of (i) the mechanisms used to condition the generative model on semantic information (e.g., text prompt) and control metadata (e.g., crop size, random flip flag, etc.) on the model performance, and (ii) the transfer of the representations learned on smaller and lower-resolution datasets to larger ones on the training efficiency and model performance. We then propose a novel conditioning mechanism that disentangles semantic and control metadata conditionings and sets a new state-of-the-art in classconditional generation on the ImageNet-1k dataset -with FID improvements of 7% on 256 and 8% on 512 resolutions -as well as text-to-image generation on the CC12M dataset -with FID improvements of 8% on 256 and 23% on 512 resolution. 38th Conference on Neural Information Processing Systems (NeurIPS 2024).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Entropy Rectifying Guidance for Diffusion and Flow ModelsTariq Berrada, Adriana Romero-Soriano, Michal Drozdzal, Jakob J. Verbeek et al.NeurIPS 2025 · 11 citations
- Flowception: Temporally Expansive Flow Matching for Video GenerationTariq Berrada Ifriqi, John Nguyen, Karteek Alahari, Jakob Verbeek et al.CVPR 2026 · 2 citations
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State SpacesKevin Rojas, Yuchen Zhu, Sichen Zhu, Felix X.-F. Ye et al.ICML 2025
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion ModelsPablo Pernias, Dominic Rampas, Mats Leon Richter, Christopher Pal et al.ICLR 2024 · 60 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Compositional Discrete Latent Code for High Fidelity, Productive Diffusion ModelsSamuel Lavoie, Michael Noukhovitch, Aaron C. CourvilleNeurIPS 2025 · 3 citations
- Guiding a Diffusion Model with a Bad Version of ItselfTero Karras, Miika Aittala, Tuomas Kynkäänniemi, Jaakko Lehtinen et al.NeurIPS 2024 · 338 citations
- Unlocking Dataset Distillation with Diffusion ModelsBrian B. Moser, Federico Raue, Sebastian Palacio, Stanislav Frolov et al.NeurIPS 2025 · 23 citations
