HL-OutPaint: Coarse-to-Fine Video Outpainting for High-Resolution Long-Range Videos
Jeongeun Park, Janghyeok Han, Geonung Kim, Hyun-Seung Lee, Kyuha Choi, Youngseok Han, Sunghyun Cho
Abstract
Video outpainting generates plausible visual content beyond a video’s original spatial extent, playing a key role in adapting videos to diverse display formats. To support such use cases, it must enable large spatial extrapolation over long sequences. However, most existing methods address only one challenge or lack explicit mechanisms, leaving notable limitations. In this paper, we propose HL-OutPaint, a high-resolution video outpainting framework for long sequences. Our approach follows a coarse-to-fine strategy with a two-stage pipeline. We first construct a Global Coarse Guidance (GCG), a low-resolution representation that captures global structure and dominant motion across the video. Unlike na"ive downsampling, the GCG is built via a novel global-local frame swapping mechanism that couples sparse global keyframes with local temporal windows and exchanges information during sampling. This enables the GCG to encode both long-term structural consistency and short-term temporal dynamics in a unified representation. Guided by this representation, HL-OutPaint then performs high-resolution outpainting to generate spatially detailed and temporally consistent content. By separating global structure modeling from fine-grained synthesis, our framework achieves stable, coherent generation for large spatial expansion and long video sequences. Extensive experiments show that HL-OutPaint outperforms existing methods in challenging scenarios with wide spatial extrapolation and long video sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9598cfa2-bf60-482f-b146-bd9dd6fba8b0Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
Related papers
- Hierarchical Masked 3D Diffusion Model for Video OutpaintingFanda Fan, Chaoxu Guo, Litong Gong, Biao Wang et al.ACM MM 2023 · 12 citations
- Infinite-Canvas: Higher-Resolution Video Outpainting with Extensive Content GenerationQihua Chen, Yue Ma, Hongfa Wang, Junkun Yuan et al.AAAI 2025 · 4 citations
- Unboxed: Geometrically and Temporally Consistent Video OutpaintingZhongrui Yu, Martina Megaro-Boldini, Robert W. Sumner, Abdelaziz DjelouahCVPR 2025
- Scenepainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentChong Xia, Shengjun Zhang, Fangfu Liu, Chang Liu et al.ICCV 2025 · 2 citations
- AnchorSync: Global Consistency Optimization for Long Video EditingZichi Liu, Yinggui Wang, Tao Wei, Chao MaACM MM 2025
