Hierarchical Patch Diffusion Models for High-Resolution Video Generation
Ivan Skorokhodov, Willi Menapace, Aliaksandr Siarohin, Sergey Tulyakov
Abstract
Diffusion models have demonstrated remarkable performance in image and video synthesis. However, scaling them to high-resolution inputs is challenging and requires restructuring the diffusion pipeline into multiple independent components, limiting scalability and complicating down-stream applications. In this work, we study patch diffusion models (PDMs) — a diffusion paradigm which models the distribution of patches, rather than whole inputs, keeping up to ≈0.7% of the original pixels. This makes it very efficient during training and unlocks end-to-end optimization on high-resolution videos. We improve PDMs in two principled ways. First, to enforce consistency between patches, we develop deep context fusion — an architectural technique that propagates the context information from low-scale to high-scale patches in a hierarchical manner. Second, to accelerate training and inference, we propose adaptive computation, which allocates more network capacity and computation towards coarse image details. The resulting model sets a new state-of-the-art FVD score of 66.32 and Inception Score of 87.68 in class-conditional video generation on UCF-101 256<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup>, surpassing recent methods by more than 100%. Then, we show that it can be rapidly fine-tuned from a base <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> low-resolution generator for high-resolution <tex xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"></tex> text-to-video synthesis. To the best of our knowledge, our model is the first diffusion-based architecture which is trained on such high resolutions entirely end-to-end. Project webpage: https://snap-research.github.io/hpdm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b73e35b-1bd9-41b3-9d7a-c4f9c0abb3b6Cited by top-tier papers8
- EVATok: Adaptive Length Video Tokenization for Efficient Visual Autoregressive GenerationTianwei Xiong, Jun Hao Liew, Zilong Huang, Zhijie Lin et al.CVPR 2026 · 8 citations
- High-Resolution Frame Interpolation with Patch-based Cascaded DiffusionJunhwa Hur, Charles Herrmann, Saurabh Saxena, Janne Kontkanen et al.AAAI 2025 · 8 citations
- Neighboring Autoregressive Modeling for Efficient Visual GenerationYefei He, Yuanyu He, Shaoxuan He, Feng Chen et al.ICCV 2025 · 3 citations
- Image Restoration via Diffusion Models with Dynamic ResolutionYang Zheng, Wen Li, Zhaoqiang LiuICML 2026 · 2 citations
- VaporTok: RL-Driven Adaptive Video Tokenizer with Prior & Task AwarenessMinghao Yang, Zechen Bai, Jing Lin, Haoqian Wang et al.NeurIPS 2025 · 1 citation
Builds on47
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Video Probabilistic Diffusion Models in Projected Latent SpaceSihyun Yu, Kihyuk Sohn, Subin Kim, Jinwoo ShinCVPR 2023
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-ResolutionShian Du, Menghan Xia, Chang Liu, Xintao Wang et al.CVPR 2025
- Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency AdapterJianhui Zhang, Sheng Cheng, Qirui Sun, Jia Liu et al.ICCV 2025 · 1 citation
- Matryoshka Diffusion ModelsJiatao Gu, Shuangfei Zhai, Yizhe Zhang, Joshua Susskind et al.ICLR 2024 · 73 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
