FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting
Rui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi, Lewei Lu, Wenxiu Sun, Xiaogang Wang, Jifeng Dai, Hongsheng Li
Abstract
Transformer, as a strong and flexible architecture for modelling long-range relations, has been widely explored in vision tasks. However, when used in video inpainting that requires fine-grained representation, existed method still suffers from yielding blurry edges in detail due to the hard patch splitting. Here we aim to tackle this problem by proposing FuseFormer, a Transformer model designed for video inpainting via fine-grained feature fusion based on novel Soft Split and Soft Composition operations. The soft split divides feature map into many patches with given overlapping interval. On the contrary, the soft composition operates by stitching different patches into a whole feature map where pixels in overlapping regions are summed up. These two modules are first used in tokenization before Transformer layers and de-tokenization after Transformer layers, for effective mapping between tokens and features. Therefore, sub-patch level information interaction is enabled for more effective feature propagation between neighboring patches, resulting in synthesizing vivid content for hole regions in videos. Moreover, in FuseFormer, we elaborately insert the soft composition and soft split into the feed-forward network, enabling the 1D linear layers to have the capability of modelling 2D structure. And, the sub-patch level feature fusion ability is further enhanced. In both quantitative and qualitative evaluations, our proposed FuseFormer surpasses state-of-the-art methods. We also conduct detailed analysis to examine its superiority. Code and pretrained models are available at https:// github.com/ruiliu-ai/FuseFormer .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e77e1fd5-4cb2-4fb6-9224-762735a45981Cited by top-tier papers51
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference TransformerZitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao et al.CVPR 2022 · 255 citations
- ViTGAN: Training GANs with Vision TransformersKwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang et al.ICLR 2022 · 225 citations
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang et al.CVPR 2022 · 214 citations
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Fast Convergence of DETR with Spatially Modulated Co-AttentionPeng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai et al.ICCV 2021 · 392 citations
Related papers
- DLFormer: Discrete Latent Transformer for Video InpaintingJingjing Ren, Qingqing Zheng, Yuanyuan Zhao, Xuemiao Xu et al.CVPR 2022 · 39 citations
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu et al.ICCV 2021 · 38 citations
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi et al.CVPR 2022 · 382 citations
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu et al.CVPR 2022 · 128 citations
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
