DeViT: Deformed Vision Transformers in Video Inpainting
Jiayin Cai, Changlin Li, Xin Tao, Chun Yuan, Yu-Wing Tai
Abstract
This paper presents a novel video inpainting architecture named Deformed Vision Transformers (DeViT). We make three significant contributions to this task: First, we extended previous Transformers with patch alignment by introducing Deformed Patch-based Homography Estimator (DePtH), which enriches the patch-level feature alignments in key and query with additional offsets learned from patch pairs without additional supervision. DePtH enables our method to handle challenging scenes or agile motion with in-plane or out-of-plane deformation, which previous methods usually fail. Second, we introduce the Mask Pruning-based Patch Attention (MPPA) to improve the standard patch-wised feature matching by pruning out less essential features and considering the saliency map. MPPA enhances the matching accuracy between warped tokens with invalid pixels. Third, we introduce the Spatial-Temporal weighting Adaptor (STA) module to assign more accurate attention to spatial-temporal tokens under the guidance of the Deformation Factor learned from DePtH, especially for videos with agile motions. Experimental results demonstrate that our method outperforms previous state-of-the-art methods in quality and quantity and achieves a new state-of-the-art for video inpainting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a74ce97-0abb-4d68-940b-d60006131a47Cited by top-tier papers3
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
- UMMAFormer: A Universal Multimodal-adaptive Transformer Framework for Temporal Forgery LocalizationRui Zhang, Hongxia Wang, Mingshan Du, Hanqing Liu et al.ACM MM 2023 · 42 citations
- DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax ConstraintsZhiliang Wu, Kun Li, Yunqiu Xu, Hehe Fan et al.AAAI 2026 · 1 citation
Builds on7
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video InpaintingRui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi et al.ICCV 2021 · 165 citations
- Copy-and-Paste Networks for Deep Video InpaintingSungho Lee, Seoung Wug Oh, DaeYeun Won, Seon Joo KimICCV 2019 · 137 citations
- Onion-Peel Networks for Deep Video CompletionSeoung Wug Oh, Sungho Lee, Joon-Young Lee, Seon Joo KimICCV 2019 · 112 citations
Related papers
- Deformable Video TransformerJue Wang, Lorenzo TorresaniCVPR 2022 · 40 citations
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi et al.CVPR 2021
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu et al.ICCV 2021 · 38 citations
- Stand-Alone Inter-Frame Attention in Video ModelsFuchen Long, Zhaofan Qiu, Yingwei Pan, Ting Yao et al.CVPR 2022 · 68 citations
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu et al.CVPR 2022 · 128 citations
