UltraGen: High-Resolution Video Generation with Hierarchical Attention
Teng Hu, Jiangning Zhang, Zihan Su, Ran Yi
2026Year
7Citations
7Top-tier citations
Abstract
UltraGen Figure 1 . Typical video generation models exhibit significant quality degradation and increased processing time with higher resolutions, whereas our UltraGen delivers superior video quality at resolutions beyond 2K while achieving 4.78× speedup compared to the popular Wan-T2V-1.3B baseline [32] (81 frames, 4×H20 GPUs). Enlarge for better visual effects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5dd59ac-eb97-4876-ac21-6f2ffa90c727Cited by top-tier papers7
- Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyTeng Hu, Zhentao Yu, Guozhen Zhang, Zihan Su et al.CVPR 2026 · 21 citations
- PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and EnhancementTeng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang et al.NeurIPS 2025 · 15 citations
- Multi-view Pyramid Transformer: Look Coarser to See BroaderGyeongjin Kang, Seungkwon Yang, Seungtae Nam, Younggeun Lee et al.CVPR 2026 · 8 citations
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal AnimationJiangning Zhang, junwei zhu, Zhenye Gan, Donghao Luo et al.CVPR 2026 · 4 citations
- Transform Trained Transformer for Accelerating Native 4K Video GenerationJiangning Zhang, Junwei Zhu, Teng Hu, Yabiao Wang et al.ICML 2026 · 3 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
Related papers
- SURF: Signature-Retained Fast Video GenerationKaixin Ding, Xi Chen, Sihui Ji, Yuan Gao et al.CVPR 2026 · 2 citations
- MAGVIT: Masked Generative Video TransformerLijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama et al.CVPR 2023
- FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisFeng Liang, Bichen Wu, Jialiang Wang, Licheng Yu et al.CVPR 2024
- Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video SynthesisJingjing Ren, Wenbo Li, Zhongdao Wang, Haoze Sun et al.ICCV 2025 · 3 citations
- SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video RestorationJianyi Wang, Zhijie Lin, Meng Wei, Yang Zhao et al.CVPR 2025
