FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
Shilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge, Peize Sun, Yifu Zhang, Yi Jiang, Zehuan Yuan, Bingyue Peng, Ping Luo
Abstract
DiT models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often require large model parameters and a substantial number of function evaluations (NFEs). Realistic and visually appealing details are typically reflected in high-resolution outputs, further amplifying computational demands—especially for single-stage DiT models. To address these challenges, we propose a novel two-stage framework, FlashVideo, which strategically allocates model capacity and NFEs across stages to balance generation fidelity and quality. In the first stage, prompt fidelity is prioritized through a low-resolution generation process utilizing large parameters and sufficient NFEs to enhance computational efficiency. The second stage achieves a nearly straight ODE trajectory between low and high resolutions via flow matching, effectively generating fine details and fixing artifacts with minimal NFEs. To ensure a seamless connection between the two independently trained stages during inference, we carefully design degradation strategies during the second-stage training. Quantitative and visual results demonstrate that FlashVideo achieves state-of-the-art high-resolution video generation with superior computational efficiency. Additionally, the two-stage design enables users to preview the initial output and accordingly adjust the prompt before committing to full-resolution generation, thereby significantly reducing computational costs and wait times as well as enhancing commercial viability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e34a3310-a0fb-4cce-bfa7-76d5220392aaCited by top-tier papers9
- Training-Free Efficient Video Generation via Dynamic Token CarvingYuechen Zhang, Jinbo Xing, Bin Xia, Shaoteng Liu et al.NeurIPS 2025 · 37 citations
- FreeViS: Training-free Video Stylization with Inconsistent ReferencesJiacong Xu, Yiqun Mei, Ke Zhang, Vishal M. PatelICLR 2026 · 7 citations
- SimpleGVR: A Simple Baseline for Latent-Cascaded Generative Video Super-ResolutionLiangbin Xie, Yu Li, Shian Du, Menghan Xia et al.ICLR 2026 · 4 citations
- Inference-time Scaling for Diffusion-based Audio Super-resolutionYizhu Jin, Zhen Ye, Zeyue Tian, Haohe Liu et al.AAAI 2026 · 3 citations
- Transform Trained Transformer for Accelerating Native 4K Video GenerationJiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang et al.ICML 2026 · 3 citations
Builds on35
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Pyramidal Flow Matching for Efficient Video Generative ModelingYang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu et al.ICLR 2025
- StreamDiT: Real-Time Streaming Text-to-Video GenerationAkio Kodaira, Tingbo Hou, Ji Hou, Markos Georgopoulos et al.CVPR 2026 · 45 citations
- Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video SynthesisJingjing Ren, Wenbo Li, Zhongdao Wang, Haoze Sun et al.ICCV 2025 · 3 citations
- LayerT2V: A Unified Multi-Layer Video Generation FrameworkGuangzhao Li, Kangrui Cen, Baixuan Zhao, Yi Xin et al.ICML 2026 · 2 citations
- REDUCIO! Generating 1K Video Within 16 Seconds Using Extremely Compressed Motion LatentsRui Tian, Qi Dai, Jianmin Bao, Kai Qiu et al.ICCV 2025 · 2 citations
