ProPainter: Improving Propagation and Transformer for Video Inpainting
Shangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change Loy
摘要
Flow-based propagation and spatiotemporal Transformer are two mainstream mechanisms in video inpainting (VI). Despite the effectiveness of these components, they still suffer from some limitations that affect their performance. Previous propagation-based approaches are performed separately either in the image or feature domain. Global image propagation isolated from learning may cause spatial misalignment due to inaccurate optical flow. Moreover, memory or computational constraints limit the temporal range of feature propagation and video Transformer, preventing exploration of correspondence information from distant frames. To address these issues, we propose an improved framework, called ProPainter, which involves enhanced ProPagation and an efficient Transformer. Specifically, we introduce dual-domain propagation that combines the advantages of image and feature warping, exploiting global correspondences reliably. We also propose a mask-guided sparse video Transformer, which achieves high efficiency by discarding unnecessary and redundant tokens. With these components, ProPainter outperforms prior arts by a large margin of 1.46 dB in PSNR while maintaining appealing efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper71
- VACE: All-in-One Video Creation and EditingZeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang 等ICCV 2025 · 被引用 58 次
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo 等CVPR 2024 · 被引用 52 次
- EasyCreator: Empowering 4D Creation through Video InpaintingYue Ma, Kunyu Feng, Xinhua Zhang, Hongyu Liu 等ICLR 2026 · 被引用 47 次
- MiniMax-Remover: Taming Bad Noise Helps Video Object RemovalBojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang 等NeurIPS 2025 · 被引用 43 次
- ROSE: Remove Objects with Side Effects in VideosChenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao 等NeurIPS 2025 · 被引用 37 次
它引用的顶会 Paper20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu 等NeurIPS 2021 · 被引用 1,343 次
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 被引用 522 次
- A-ViT: Adaptive Tokens for Efficient Vision TransformerHongxu Yin, Arash Vahdat, José M. Álvarez, Arun Mallya 等CVPR 2022 · 被引用 288 次
- Focal Attention for Long-Range Interactions in Vision TransformersJianwei Yang, Chunyuan Li, Pengchuan Zhang, Xiyang Dai 等NeurIPS 2021 · 被引用 228 次
相关 Paper
- Progressive Temporal Feature Alignment Network for Video InpaintingXueyan Zou, Linjie Yang, Ding Liu, Yong Jae LeeCVPR 2021
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 等ICCV 2021 · 被引用 38 次
- Blur-Aware Spatio-Temporal Sparse Transformer for Video DeblurringHuicong Zhang, Haozhe Xie, Hongxun YaoCVPR 2024 · 被引用 14 次
- Video Frame Interpolation with Flow TransformerPan Gao, Haoyue Tian, Jie QinACM MM 2023 · 被引用 4 次
- VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context ControlYuxuan Bian, Zhaoyang Zhang, Xuan Ju, Mingdeng Cao 等SIGGRAPH 2025 · 被引用 11 次
