Frequency-Aware Spatiotemporal Transformers for Video Inpainting Detection
Bingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu, Jie Zhou
Abstract
In this paper, we propose a Frequency-Aware Spatiotemporal Transformer (FAST) for video inpainting detection, which aims to simultaneously mine the traces of video in-painting from spatial, temporal, and frequency domains. Unlike existing deep video inpainting detection methods that usually rely on hand-designed attention modules and memory mechanism, our proposed FAST have innate global self-attention mechanisms to capture the long-range relations. While existing video inpainting methods usually exploit the spatial and temporal connections in a video, our method employs a spatiotemporal transformer framework to detect the spatial connections between patches and temporal dependency between frames. As the inpainted videos usually lack high frequency details, our proposed FAST synchronously exploits the frequency domain information with a specifically designed decoder. Extensive experimental results demonstrate that our approach achieves very competitive performance and generalizes well.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9a9c6f8-fac2-45aa-9102-4f0961552974Builds on12
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Attention Augmented Convolutional NetworksIrwan Bello, Barret Zoph, Quoc Le, Ashish Vaswani et al.ICCV 2019 · 1,149 citations
- Leveraging Frequency Analysis for Deep Fake Image RecognitionJoel Frank, Thorsten Eisenhofer, Lea Schönherr, Asja Fischer et al.ICML 2020 · 848 citations
- Local Relation Networks for Image RecognitionHan Hu, Zheng Zhang, Zhenda Xie, Stephen LinICCV 2019 · 555 citations
- Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGANYa-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston H. HsuICCV 2019 · 213 citations
Related papers
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
- Beyond Spatial Frequency: Pixel-Wise Temporal Frequency-Based Deepfake Video DetectionTaehoon Kim, Jongwook Choi, Yonghyun Jeong, Haeun Noh et al.ICCV 2025 · 7 citations
- Learning Contextual Transformer Network for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng et al.ACM MM 2021 · 29 citations
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
