DLFormer: Discrete Latent Transformer for Video Inpainting
Jingjing Ren, Qingqing Zheng, Yuanyuan Zhao, Xuemiao Xu, Chen Li
Abstract
Video inpainting remains a challenging problem to fill with plausible and coherent content in unknown areas in video frames despite the prevalence of data-driven methods. Although various transformer-based architectures yield promising result for this task, they still suffer from hallucinating blurry contents and long-term spatial-temporal inconsistency. While noticing the capability of discrete representation for complex reasoning and predictive learning, we propose a novel Discrete Latent Transformer (DLFormer) to reformulate video inpainting tasks into the discrete latent space rather the previous continuous feature space. Specifically, we first learn a unique compact discrete codebook and the corresponding autoencoder to represent the target video. Built upon these representative discrete codes obtained from the entire target video, the subsequent discrete latent transformer is capable to infer proper codes for unknown areas under a self-attention mechanism, and thus produces fine-grained content with long-term spatial-temporal consistency. Moreover, we further explicitly enforce the short-term consistency to relieve temporal visual jitters via a temporal aggregation block among adjacent frames. We conduct comprehensive quantitative and qualitative evaluations to demonstrate that our method significantly outperforms other state-of-the-art approaches in reconstructing visually-plausible and spatial-temporal coherent content with fine-grained details. Code is available at https://github.com/JingjingRenabc/dlformer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0adb622f-2a4a-4ef6-97cf-e524221edd21Cited by top-tier papers17
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 205 citations
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu et al.AAAI 2024 · 85 citations
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 70 citations
- Adverse Weather Removal with Codebook PriorsTian Ye, Sixiang Chen, Jinbin Bai, Jun Shi et al.ICCV 2023 · 64 citations
- Snow Removal in Video: A New Dataset and A Novel MethodHaoyu Chen, Jingjing Ren, Jinjin Gu, Hongtao Wu et al.ICCV 2023 · 40 citations
Builds on10
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGANYa-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston H. HsuICCV 2019 · 213 citations
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video InpaintingRui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi et al.ICCV 2021 · 165 citations
- Copy-and-Paste Networks for Deep Video InpaintingSungho Lee, Seoung Wug Oh, DaeYeun Won, Seon Joo KimICCV 2019 · 137 citations
- Onion-Peel Networks for Deep Video CompletionSeoung Wug Oh, Sungho Lee, Joon-Young Lee, Seon Joo KimICCV 2019 · 112 citations
Related papers
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen et al.CVPR 2022 · 117 citations
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu et al.ICCV 2021 · 38 citations
- Don't Look into the Dark: Latent Codes for Pluralistic Image InpaintingHaiwei Chen, Yajie ZhaoCVPR 2024
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu et al.CVPR 2022 · 128 citations
- Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video RestorationLianxin Xie, Bingbing Zheng, Wen Xue, Yunfei Zhang et al.AAAI 2025
