Infinite-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
Qihua Chen, Yue Ma, Hongfa Wang, Junkun Yuan, Wenzhe Zhao, Qi Tian, Hongmei Wang, Shaobo Min, Qifeng Chen, Wei Liu
Abstract
This paper explores higher-resolution video outpainting with extensive content generation. We point out common issues faced by existing methods when attempting to largely outpaint videos: the generation of low-quality content and limitations imposed by GPU memory. To address these challenges, we propose a diffusion-based method called Infinite-Canvas. It builds upon two core designs. First, instead of employing the common practice of "single-shot" outpainting, we distribute the task across spatial windows and seamlessly merge them. It allows us to outpaint videos of any size and resolution without being constrained by GPU memory. Second, the source video and its relative positional relation are injected into the generation process of each window. It makes the generated spatial layout within each window harmonize with the source video. Coupling with these two designs enables us to generate higher-resolution outpainting videos with rich content while keeping spatial and temporal consistency. Infinite-Canvas excels in large-scale video outpainting, e.g., from 512 × 512 to 1152 × 2048 (9×), while producing high-quality and aesthetically pleasing results. It achieves the best quantitative results across various resolution and scale setups. The code is available at https://github.com/mayuelala/FollowYourCanvas.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range VideosJeongeun Park, Janghyeok Han, Geonung Kim, Hyun-Seung Lee et al.SIGGRAPH 2026
- Tea-Adapter: Teacher Adapter for Efficient Conditional GenerationYinhan Zhang, Yue Ma, Fangqiu Yi, Chenyang Qi et al.CVPR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Hierarchical Masked 3D Diffusion Model for Video OutpaintingFanda Fan, Chaoxu Guo, Litong Gong, Biao Wang et al.ACM MM 2023 · 12 citations
- InfinityGAN: Towards Infinite-Pixel Image SynthesisChieh Hubert Lin, Hsin-Ying Lee, Yen-Chi Cheng, Sergey Tulyakov et al.ICLR 2022 · 84 citations
- Unboxed: Geometrically and Temporally Consistent Video OutpaintingZhongrui Yu, Martina Megaro-Boldini, Robert W. Sumner, Abdelaziz DjelouahCVPR 2025
- Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based ApproachShaofeng Zhang, Jinfa Huang, Qiang Zhou, Zhibin Wang et al.ICLR 2024 · 24 citations
- VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context ControlYuxuan Bian, Zhaoyang Zhang, Xuan Ju, Mingdeng Cao et al.SIGGRAPH 2025 · 11 citations
