PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient Inference
Denis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian, Amirhossein Habibian
Abstract
Recently proposed pyramidal models decompose the conventional forward and backward diffusion processes into multiple stages operating at varying resolutions. These models handle inputs with higher noise levels at lower resolutions, while less noisy inputs are processed at higher resolutions. This hierarchical approach significantly reduces the computational cost of inference in multi-step denoising models. However, existing open-source pyramidal video models have been trained from scratch and tend to underperform compared to state-of-the-art systems in terms of visual plausibility. In this work, we present a pipeline that converts a pretrained diffusion model into a pyramidal one through low-cost finetuning, achieving this transformation without degradation in quality of output videos. Furthermore, we investigate and compare various strategies for step distillation within pyramidal models, aiming to further enhance the inference efficiency. Our results are available at https://qualcomm-ai-research.github. io/PyramidalWan * Equal contribution † Qualcomm AI Research is an initiative of Qualcomm Technologies, Inc. Snapdragon and Qualcomm branded products are products of Qualcomm Technologies, Inc. and/or its subsidiaries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e1e62e9a-12a7-4703-bffb-68a15e712d21Builds on35
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang et al.NeurIPS 2024 · 728 citations
- Mean Flows for One-step Generative ModelingZhengyang Geng, Mingyang Deng, Xingjian Bai, Zico Kolter et al.NeurIPS 2025 · 628 citations
- simple diffusion: End-to-end diffusion for high resolution imagesEmiel Hoogeboom, Jonathan Heek, Tim SalimansICML 2023 · 403 citations
Related papers
- Mobile Video DiffusionHaitam Ben Yahia, Denis Korzhenkov, Ioannis Lelekas, Amir Ghodrati et al.ICCV 2025 · 1 citation
- Improving Progressive Generation with Decomposable Flow MatchingMoayed Haji-Ali, Willi Menapace, Ivan Skorokhodov, Arpit Sahni et al.NeurIPS 2025 · 7 citations
- Streaming Autoregressive Video Generation via Diagonal DistillationJinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen et al.ICLR 2026 · 16 citations
- Pyramidal Flow Matching for Efficient Video Generative ModelingYang Jin, Zhicheng Sun, Ningyuan Li, Kun Xu et al.ICLR 2025
- Minute-Long Videos with Dual ParallelismsZeqing Wang, Bowen Zheng, Xingyi Yang, Zhenxiong Tan et al.AAAI 2026
