Deeply Supervised Flow-Based Generative Models
Inkyu Shin, Chenglin Yang, Liang-Chieh Chen
摘要
Flow-based generative models have charted an impressive path across multiple visual generation tasks by adhering to a simple principle: learning velocity representations of a linear interpolant. However, we observe that training velocity solely from the final layer's output under-utilizes the rich inter-layer representations, potentially impeding model convergence. To address this limitation, we introduce DeepFlow, a novel framework that enhances velocity representation through inter-layer communication. DeepFlow partitions transformer layers into balanced branches with deep supervision and inserts a lightweight Velocity Refiner with Acceleration (VeRA) block between adjacent branches, which aligns the intermediate velocity features within transformer blocks. Powered by the improved deep supervision via the internal velocity alignment, DeepFlow converges faster on ImageNet-256 with equivalent performance and further reduces FID by while halving training time compared to previous flow-based models without a classifier-free guidance. DeepFlow also outperforms baselines in text-to-image generation tasks, as evidenced by evaluations on MS-COCO and zero-shot GenEval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Guiding a Diffusion Transformer with the Internal Dynamics of ItselfXingyu Zhou, Qifan Li, Xiaobin Hu, Hai Chen 等CVPR 2026 · 被引用 13 次
- FlowTok: Flowing Seamlessly Across Text and Image TokensJu He, Qihang Yu, Qihao Liu, Liang-Chieh ChenICCV 2025 · 被引用 8 次
- A Frame is Worth One Token: Efficient Generative World Modeling with Delta TokensTommie Kerssies, Gabriele Berton, Ju He, Qihang Yu 等CVPR 2026 · 被引用 8 次
- Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional TokensDongwon Kim, Ju He, Qihang Yu, Chenglin Yang 等ICCV 2025 · 被引用 8 次
- Frequency-Aware Flow Matching for High-Quality Image GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen 等CVPR 2026 · 被引用 6 次
它引用的顶会 Paper28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity RefinerWenliang Zhao, Minglei Shi, Xumin Yu, Jie Zhou 等NeurIPS 2024 · 被引用 6 次
- VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and EstimationJunwen Tan, Jinglin Liang, Hongyuan Chen, Shuangping HuangCVPR 2026 · 被引用 1 次
- Stable Velocity: A Variance Perspective on Flow MatchingDonglin Yang, Yongxing Zhang, Xin Yu, Liang Hou 等ICML 2026 · 被引用 6 次
- Straighten Viscous Rectified Flow via Noise OptimizationJimin Dai, Jiexi Yan, Jian Yang, Lei LuoICCV 2025 · 被引用 1 次
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality GenerationDogyun Park, Taehoon Lee, Minseok Joo, Hyunwoo J. KimNeurIPS 2025 · 被引用 4 次
