Pyramidal Flow Matching for Efficient Video Generative Modeling
Yang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu, Hao Jiang, Nan Zhuang, Quzhe Huang, Yang Song, Yadong Mu, Zhouchen Lin
摘要
Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training with full resolution latent. Despite reducing computational demands, the separate optimization of each sub-stage hinders knowledge sharing and sacrifices flexibility. This work introduces a unified pyramidal flow matching algorithm. It reinterprets the original denoising trajectory as a series of pyramid stages, where only the final stage operates at the full resolution, thereby enabling more efficient video generative modeling. Through our sophisticated design, the flows of different pyramid stages can be interlinked to maintain continuity. Moreover, we craft autoregressive video generation with a temporal pyramid to compress the full-resolution history. The entire framework can be optimized in an end-to-end manner and with a single unified Diffusion Transformer (DiT). Extensive experiments demonstrate that our method supports generating high-quality 5-second (up to 10-second) videos at 768p resolution and 24 FPS within 20.7k A100 GPU training hours. All code and models are open-sourced at https://pyramid-flow.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper147
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao 等ICLR 2026 · 被引用 241 次
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan 等ICLR 2026 · 被引用 215 次
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang 等ICLR 2026 · 被引用 181 次
- WorldMem: Long-term Consistent World Simulation with MemoryZeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang 等NeurIPS 2025 · 被引用 165 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin 等ICCV 2025
- Improving Progressive Generation with Decomposable Flow MatchingMoayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni 等NeurIPS 2025 · 被引用 7 次
- FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video GenerationShilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge 等AAAI 2026 · 被引用 30 次
- Real-Time Video Generation with Pyramid Attention BroadcastXuanlei Zhao, Xiaolong Jin, Kai Wang, Yang YouICLR 2025
- Pyramid Patchification Flow for Visual GenerationHui Li, Baoyou Chen, Jiaye Li, Jingdong Wang 等ICLR 2026 · 被引用 1 次
