PaGoDA: Progressive Growing of a One-Step Generator from a Low-Resolution Diffusion Teacher
Dongjun Kim, Chieh-Hsin Lai, Wei-Hsiang Liao, Yuhta Takida, Naoki Murata, Toshimitsu Uesaka, Yuki Mitsufuji, Stefano Ermon
摘要
The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the training costs through three stages: training diffusion on downsampled data, distilling the pretrained diffusion, and progressive super-resolution. With the proposed pipeline, PaGoDA achieves a reduced cost in training its diffusion model on 8x downsampled data; while at the inference, with the single-step, it performs state-of-the-art on ImageNet across all resolutions from 64x64 to 512x512, and text-to-image. PaGoDA's pipeline can be applied directly in the latent space, adding compression alongside the pre-trained autoencoder in Latent Diffusion Models (e.g., Stable Diffusion). The code is available at https://github.com/sony/pagoda.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion ModelsShuchen Xue, Chongjian GE, Shilong Zhang, Yichen Li 等ICML 2026 · 被引用 43 次
- Scale-wise Distillation of Diffusion ModelsNikita Starodubcev, Ilya Drobyshevskiy, Denis Kuznedelev, Artem Babenko 等ICLR 2026 · 被引用 13 次
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationJunsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu 等ICCV 2025 · 被引用 6 次
- MeanFlow Transformers with Representation AutoencodersZheyuan Hu, Chieh-Hsin Lai, Ge Wu, Yuki Mitsufuji 等CVPR 2026 · 被引用 6 次
- Latent Stochastic InterpolantsSaurabh Singh, Dmitry LagunICLR 2026 · 被引用 2 次
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image GenerationClément Chadebec, Onur Tasar, Eyal Benaroche, Benjamin AubinAAAI 2025 · 被引用 52 次
- Progressive Distillation for Fast Sampling of Diffusion ModelsTim Salimans, Jonathan HoICLR 2022 · 被引用 9 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Würstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion ModelsPablo Pernias, Dominic Rampas, Mats Leon Richter, Christopher Pal 等ICLR 2024 · 被引用 60 次
- DC-AE 1.5: Accelerating Diffusion Model Convergence with Structured Latent SpaceJunyu Chen, Dongyun Zou, Wenkun He, Junsong Chen 等ICCV 2025 · 被引用 3 次
