f-DM: A Multi-stage Diffusion Model via Progressive Signal Transformation
Jiatao Gu, Shuangfei Zhai, Yizhe Zhang, Miguel Ángel Bautista, Joshua M. Susskind
摘要
Diffusion models (DMs) have recently emerged as SoTA tools for generative modeling in various domains. Standard DMs can be viewed as an instantiation of hierarchical variational autoencoders (VAEs) where the latent variables are inferred from input-centered Gaussian distributions with fixed scales and variances. Unlike VAEs, this formulation constrains DMs from changing the latent spaces and learning abstract representations. In this work, we propose f -DM, a generalized family of DMs, which allows progressive signal transformation. More precisely, we extend DMs to incorporate a set of (hand-designed or learned) transformations, where the transformed input is the mean of each diffusion step. We propose a generalized formulation of DMs and derive the corresponding de-noising objective together with a modified sampling algorithm. As a demonstration, we apply f -DM in image generation tasks with a range of functions, including down-sampling, blurring, and learned transformations based on the encoder of pretrained VAEs. In addition, we identify the importance of adjusting the noise levels whenever the signal is sub-sampled and propose a simple rescaling recipe. f -DM can produce high-quality samples on standard image generation benchmarks like FFHQ, AFHQ, LSUN and ImageNet with better efficiency and semantic interpretation. Please check our videos at http://jiataogu.me/fdm/ . Figure 1 : Visualization of reverse diffusion from f -DMs with various signal transformations. x t is the denoised output, and z s is the input to the next diffusion step. We plot the first three channels of VQVAE latent variables. Low-resolution images are resized to 256 2 for ease of visualization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- FreCaS: Efficient Higher-Resolution Image Generation via Frequency-aware Cascaded SamplingZhengqiang Zhang, Ruihuang Li, Lei ZhangICLR 2025
- Simpler Diffusion: 1.5 FID on ImageNet512 with Pixel-space DiffusionEmiel Hoogeboom, Thomas Mensink, Jonathan Heek, Kay Lamerigts 等CVPR 2025
- FM-Steer: Enhance Generalist Policies with Value-Guided Cascaded DenoisingHaoming Song, Delin Qu, Yuanqi Yao, Qizhi Chen 等CVPR 2026
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Neural Diffusion ModelsGrigory Bartosh, Dmitry P. Vetrov, Christian A. NaessethICML 2024 · 被引用 69 次
- Progressive Distillation for Fast Sampling of Diffusion ModelsTim Salimans, Jonathan HoICLR 2022 · 被引用 9 次
- UDPM: Upsampling Diffusion Probabilistic ModelsShady Abu-Hussein, Raja GiryesNeurIPS 2024 · 被引用 10 次
- Learning Fast Samplers for Diffusion Models by Differentiating Through Sample QualityDaniel Watson, William Chan, Jonathan Ho, Mohammad NorouziICLR 2022 · 被引用 224 次
- Improving Progressive Generation with Decomposable Flow MatchingMoayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni 等NeurIPS 2025 · 被引用 7 次
