Improving Progressive Generation with Decomposable Flow Matching
Moayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni, Sergey Tulyakov, Vicente Ordonez, Aliaksandr Siarohin
摘要
Generating high-dimensional visual modalities is a computationally intensive task. A common solution is progressive generation, where the outputs are synthesized in a coarse-to-fine spectral autoregressive manner. While diffusion models benefit from the coarse-to-fine nature of denoising, explicit multi-stage architectures are rarely adopted. These architectures have increased the complexity of the overall approach, introducing the need for a custom diffusion formulation, decomposition-dependent stage transitions, add-hoc samplers, or a model cascade. Our contribution, Decomposable Flow Matching (DFM), is a simple and effective framework for the progressive generation of visual media. DFM applies Flow Matching independently at each level of a user-defined multi-scale representation (such as Laplacian pyramid). As shown by our experiments, our approach improves visual quality for both images and videos, featuring superior results compared to prior multistage frameworks. On Imagenet-1k 512px, DFM achieves 35.2% improvements in FDD scores over the base architecture and 26.4% over the best-performing baseline, under the same training compute. When applied to finetuning of large models, such as FLUX, DFM shows faster convergence speed to the training distribution. Crucially, all these advantages are achieved with a single model, architectural simplicity, and minimal modifications to existing training pipelines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Scale-wise Distillation of Diffusion ModelsNikita Starodubcev, Ilya Drobyshevskiy, Denis Kuznedelev, Artem Babenko 等ICLR 2026 · 被引用 13 次
- One Model, Many Budgets: Elastic Latent Interfaces for Diffusion TransformersMoayed Haji Ali, Willi Menapace, Ivan Skorokhodov, Dogyun Park 等CVPR 2026 · 被引用 4 次
- Scale Space DiffusionSoumik Mukhopadhyay, Prateksha Udhayanan, Abhinav ShrivastavaCVPR 2026 · 被引用 4 次
它引用的顶会 Paper28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- LapFlow: Laplacian Multi-scale Flow Matching for Generative ModelingZelin Zhao, Petr Molodyk, Haotian Xue, Yongxin ChenICLR 2026 · 被引用 2 次
- Pyramidal Flow Matching for Efficient Video Generative ModelingYang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu 等ICLR 2025
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman 等CVPR 2024 · 被引用 75 次
- Terminal Velocity MatchingLinqi Zhou, Mathias Parger, Ayaan Haque, Jiaming SongICLR 2026 · 被引用 15 次
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu 等CVPR 2026 · 被引用 24 次
