Deeply Supervised Flow-Based Generative Models
Inkyu Shin, Chenglin Yang, Liang-Chieh Chen
Abstract
Flow-based generative models have charted an impressive path across multiple visual generation tasks by adhering to a simple principle: learning velocity representations of a linear interpolant. However, we observe that training velocity solely from the final layer's output under-utilizes the rich inter-layer representations, potentially impeding model convergence. To address this limitation, we introduce DeepFlow, a novel framework that enhances velocity representation through inter-layer communication. DeepFlow partitions transformer layers into balanced branches with deep supervision and inserts a lightweight Velocity Refiner with Acceleration (VeRA) block between adjacent branches, which aligns the intermediate velocity features within transformer blocks. Powered by the improved deep supervision via the internal velocity alignment, DeepFlow converges faster on ImageNet-256 with equivalent performance and further reduces FID by while halving training time compared to previous flow-based models without a classifier-free guidance. DeepFlow also outperforms baselines in text-to-image generation tasks, as evidenced by evaluations on MS-COCO and zero-shot GenEval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Guiding a Diffusion Transformer with the Internal Dynamics of ItselfXingyu Zhou, Qifan Li, Xiaobin Hu, Hai Chen et al.CVPR 2026 · 13 citations
- FlowTok: Flowing Seamlessly Across Text and Image TokensJu He, Qihang Yu, Qihao Liu, Liang-Chieh ChenICCV 2025 · 8 citations
- A Frame is Worth One Token: Efficient Generative World Modeling with Delta TokensTommie Kerssies, Gabriele Berton, Ju He, Qihang Yu et al.CVPR 2026 · 8 citations
- Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional TokensDongwon Kim, Ju He, Qihang Yu, Chenglin Yang et al.ICCV 2025 · 8 citations
- Frequency-Aware Flow Matching for High-Quality Image GenerationSucheng Ren, Qihang Yu, Ju He, Xiaohui Shen et al.CVPR 2026 · 6 citations
Builds on28
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- FlowTurbo: Towards Real-time Flow-Based Image Generation with Velocity RefinerWenliang Zhao, Minglei Shi, Xumin Yu, Jie Zhou et al.NeurIPS 2024 · 6 citations
- VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and EstimationJunwen Tan, Jinglin Liang, Hongyuan Chen, Shuangping HuangCVPR 2026 · 1 citation
- Stable Velocity: A Variance Perspective on Flow MatchingDonglin Yang, Yongxing Zhang, Xin Yu, Liang Hou et al.ICML 2026 · 6 citations
- Straighten Viscous Rectified Flow via Noise OptimizationJimin Dai, Jiexi Yan, Jian Yang, Lei LuoICCV 2025 · 1 citation
- Blockwise Flow Matching: Improving Flow Matching Models For Efficient High-Quality GenerationDogyun Park, Taehoon Lee, Minseok Joo, Hyunwoo J. KimNeurIPS 2025 · 4 citations
