Occlusion-Aware Video Object Inpainting
Lei Ke, Yu-Wing Tai, Chi-Keung Tang
Abstract
Conventional video inpainting is neither object-oriented nor occlusion-aware, making it liable to obvious artifacts when large occluded object regions are inpainted. This paper presents occlusion-aware video object inpainting, which recovers both the complete shape and appearance for occluded objects in videos given their visible mask segmentation. To facilitate this new research, we construct the first large-scale video object inpainting benchmark YouTube-VOI to provide realistic occlusion scenarios with both occluded and visible object masks available. Our technical contribution VOIN jointly performs video object shape completion and occluded texture generation. In particular, the shape completion module models long-range object coherence while the flow completion module recovers accurate flow with sharp motion boundary, for propagating temporally-consistent texture to the same moving object across frames. For more realistic results, VOIN is optimized using both T-PatchGAN and a new spatio-temporal attention-based multi-class discriminator. Finally, we compare VOIN and strong baselines on YouTube-VOI. Experimental results clearly demonstrate the efficacy of our method including inpainting complex and dynamic objects. VOIN degrades gracefully with inaccurate input visible mask.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ae8ba4c-1602-4414-8d8a-9381be6750d2Cited by top-tier papers16
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
- Prototypical Cross-Attention Networks for Multiple Object Tracking and SegmentationLei Ke, Xia Li, Martin Danelljan, Yu-Wing Tai et al.NeurIPS 2021 · 92 citations
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
- Inertia-Guided Flow Completion and Style Fusion for Video InpaintingKaidong Zhang, Jingjing Fu, Dong LiuCVPR 2022 · 42 citations
- Coarse-to-Fine Amodal Segmentation with Shape PriorJianxiong Gao, Xuelin Qian, Yikai Wang, Tianjun Xiao et al.ICCV 2023 · 36 citations
Builds on14
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Video Instance SegmentationLinjie Yang, Yuchen Fan, Ning XuICCV 2019 · 615 citations
- Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGANYa-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston H. HsuICCV 2019 · 213 citations
- Copy-and-Paste Networks for Deep Video InpaintingSungho Lee, Seoung Wug Oh, DaeYeun Won, Seon Joo KimICCV 2019 · 137 citations
Related papers
- Dynamic Shadow Unveils Invisible Semantics for Video OutpaintingRuilin Li, Hang Yu, Jiayan QiuNeurIPS 2025 · 2 citations
- Using Diffusion Priors for Video Amodal SegmentationKaihua Chen, Deva Ramanan, Tarasha KhuranaCVPR 2025
- Elevating Flow-Guided Video Inpainting with Reference GenerationSuhwan Cho, Seoung Wug Oh, Sangyoun Lee, Joon-Young LeeAAAI 2025 · 2 citations
- An Internal Learning Approach to Video InpaintingHaotian Zhang, Long Mai, Hailin Jin, Zhaowen Wang et al.ICCV 2019 · 77 citations
- Flow-Guided Video Inpainting with Scene TemplatesDong Lao, Peihao Zhu, Peter Wonka, Ganesh SundaramoorthiICCV 2021 · 18 citations
