DLFormer: Discrete Latent Transformer for Video Inpainting
Jingjing Ren, Qingqing Zheng, Yuanyuan Zhao, Xuemiao Xu, Chen Li
摘要
Video inpainting remains a challenging problem to fill with plausible and coherent content in unknown areas in video frames despite the prevalence of data-driven methods. Although various transformer-based architectures yield promising result for this task, they still suffer from hallucinating blurry contents and long-term spatial-temporal inconsistency. While noticing the capability of discrete representation for complex reasoning and predictive learning, we propose a novel Discrete Latent Transformer (DLFormer) to reformulate video inpainting tasks into the discrete latent space rather the previous continuous feature space. Specifically, we first learn a unique compact discrete codebook and the corresponding autoencoder to represent the target video. Built upon these representative discrete codes obtained from the entire target video, the subsequent discrete latent transformer is capable to infer proper codes for unknown areas under a self-attention mechanism, and thus produces fine-grained content with long-term spatial-temporal consistency. Moreover, we further explicitly enforce the short-term consistency to relieve temporal visual jitters via a temporal aggregation block among adjacent frames. We conduct comprehensive quantitative and qualitative evaluations to demonstrate that our method significantly outperforms other state-of-the-art approaches in reconstructing visually-plausible and spatial-temporal coherent content with fine-grained details. Code is available at https://github.com/JingjingRenabc/dlformer.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 被引用 205 次
- WaveFormer: Wavelet Transformer for Noise-Robust Video InpaintingZhiliang Wu, Changchang Sun, Hanyu Xuan, Gaowen Liu 等AAAI 2024 · 被引用 85 次
- SG-Former: Self-guided Transformer with Evolving Token ReallocationSucheng Ren, Xingyi Yang, Songhua Liu, Xinchao WangICCV 2023 · 被引用 70 次
- Adverse Weather Removal with Codebook PriorsTian Ye, Sixiang Chen, Jinbin Bai, Jun Shi 等ICCV 2023 · 被引用 64 次
- Snow Removal in Video: A New Dataset and A Novel MethodHaoyu Chen, Jingjing Ren, Jinjin Gu, Hongtao Wu 等ICCV 2023 · 被引用 40 次
它引用的顶会 Paper10
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 被引用 2,049 次
- Free-Form Video Inpainting With 3D Gated Convolution and Temporal PatchGANYa-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, Winston H. HsuICCV 2019 · 被引用 213 次
- FuseFormer: Fusing Fine-Grained Information in Transformers for Video InpaintingRui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi 等ICCV 2021 · 被引用 165 次
- Copy-and-Paste Networks for Deep Video InpaintingSungho Lee, Seoung Wug Oh, DaeYeun Won, Seon Joo KimICCV 2019 · 被引用 137 次
- Onion-Peel Networks for Deep Video CompletionSeoung Wug Oh, Sungho Lee, Joon-Young Lee, Seon Joo KimICCV 2019 · 被引用 112 次
相关 Paper
- Video Frame Interpolation TransformerZhihao Shi, Xiangyu Xu, Xiaohong Liu, Jun Chen 等CVPR 2022 · 被引用 117 次
- Frequency-Aware Spatiotemporal Transformers for Video Inpainting DetectionBingyao Yu, Wanhua Li, Xiu Li, Jiwen Lu 等ICCV 2021 · 被引用 38 次
- Don't Look into the Dark: Latent Codes for Pluralistic Image InpaintingHaiwei Chen, Yajie ZhaoCVPR 2024
- Video Frame Interpolation with TransformerLiying Lu, Ruizheng Wu, Huaijia Lin, Jiangbo Lu 等CVPR 2022 · 被引用 128 次
- Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video RestorationLianxin Xie, Bingbing Zheng, Wen Xue, Yunfei Zhang 等AAAI 2025
