FuseFormer: Fusing Fine-Grained Information in Transformers for Video Inpainting
Rui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi, Lewei Lu, Wenxiu Sun, Xiaogang Wang, Jifeng Dai, Hongsheng Li
摘要
Transformer, as a strong and flexible architecture for modelling long-range relations, has been widely explored in vision tasks. However, when used in video inpainting that requires fine-grained representation, existed method still suffers from yielding blurry edges in detail due to the hard patch splitting. Here we aim to tackle this problem by proposing FuseFormer, a Transformer model designed for video inpainting via fine-grained feature fusion based on novel Soft Split and Soft Composition operations. The soft split divides feature map into many patches with given overlapping interval. On the contrary, the soft composition operates by stitching different patches into a whole feature map where pixels in overlapping regions are summed up. These two modules are first used in tokenization before Transformer layers and de-tokenization after Transformer layers, for effective mapping between tokens and features. Therefore, sub-patch level information interaction is enabled for more effective feature propagation between neighboring patches, resulting in synthesizing vivid content for hole regions in videos. Moreover, in FuseFormer, we elaborately insert the soft composition and soft split into the feed-forward network, enabling the 1D linear layers to have the capability of modelling 2D structure. And, the sub-patch level feature fusion ability is further enhanced. In both quantitative and qualitative evaluations, our proposed FuseFormer surpasses state-of-the-art methods. We also conduct detailed analysis to examine its superiority. Code and pretrained models are available at https:// github.com/ruiliu-ai/FuseFormer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper51
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu 等NeurIPS 2022 · 被引用 1,385 次
- PhysFormer: Facial Video-based Physiological Measurement with Temporal Difference TransformerZitong Yu, Yuming Shen, Jingang Shi, Hengshuang Zhao 等CVPR 2022 · 被引用 255 次
- ViTGAN: Training GANs with Vision TransformersKwonjoon Lee, Huiwen Chang, Lu Jiang, Han Zhang 等ICLR 2022 · 被引用 225 次
- Multimodal Token Fusion for Vision TransformersYikai Wang, Xinghao Chen, Lele Cao, Wenbing Huang 等CVPR 2022 · 被引用 214 次
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 被引用 205 次
它引用的顶会 Paper11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen 等ICCV 2019 · 被引用 1,990 次
- Fast Convergence of DETR with Spatially Modulated Co-AttentionPeng Gao, Minghang Zheng, Xiaogang Wang, Jifeng Dai 等ICCV 2021 · 被引用 392 次
相关 Paper
- DLFormer: Discrete Latent Transformer for Video InpaintingJingjing Ren, Qingqing Zheng, Yuanyuan Zhao, Xuemiao Xu 等CVPR 2022 · 被引用 39 次
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 等ICCV 2021 · 被引用 38 次
- MAT: Mask-Aware Transformer for Large Hole Image InpaintingWenbo Li, Zhe Lin, Kun Zhou, Lu Qi 等CVPR 2022 · 被引用 382 次
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu 等CVPR 2022 · 被引用 128 次
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu 等AAAI 2024 · 被引用 85 次
