VRetouchEr: Learning Cross-Frame Feature Interdependence with Imperfection Flow for Face Retouching in Videos
Wen Xue, Le Jiang, Lianxin Xie, Si Wu, Yong Xu, Hau-San Wong
Abstract
Face Video Retouching is a complex task that often requires labor-intensive manual editing. Conventional image retouching methods perform less satisfactorily in terms of generalization performance and stability when applied to videos without exploiting the correlation among frames. To address this issue, we propose a Video Retouching transformEr to remove facial imperfections in videos, which is referred to as VRetouchEr. Specifically, we estimate the apparent motion of imperfections between two consecutive frames, and the resulting displacement vectors are used to refine the imperfection map, which is synthesized from the current frame together with the corresponding encoder features. The flow-based imperfection refinement is critical for precise and stable retouching across frames. To leverage the temporal contextual information, we inject the refined imperfection map into each transformer block for multi-frame masked attention computation, such that we can capture the interdependence between the current frame and multiple reference frames. As a result, the imperfection regions can be replaced with normal skin with high fidelity, while at the same time keeping the other regions unchanged. Extensive experiments are performed to verify the superiority of VRetouchEr over state-of-the-art image retouching methods in terms of fidelity and stability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72fff991-3fd7-42e6-8461-73e14ae451e3Cited by top-tier papers2
- RetouchGPT: LLM-based Interactive High-Fidelity Face Retouching via Imperfection PromptingWen Xue, Chun Ding, Ruotao Xu, Si Wu et al.AAAI 2025 · 3 citations
- BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference ModelingJiachen Yang, Xianhui Lin, Yi Dong, Zebiao Zheng et al.CVPR 2026 · 1 citation
Builds on26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- On the "steerability" of generative adversarial networksAli Jahanian, Lucy Chai, Phillip IsolaICLR 2020 · 421 citations
- EditGAN: High-Precision Semantic Image EditingHuan Ling, Karsten Kreis, Daiqing Li, Seung Wook Kim et al.NeurIPS 2021 · 248 citations
Related papers
- RetouchFormer: Semi-supervised High-Quality Face Retouching Transformer with Prior-Based Selective Self-AttentionXue Wen, Lianxin Xie, Le Jiang, Tianyi Chen et al.AAAI 2024 · 3 citations
- Hunting Blemishes: Language-guided High-fidelity Face Retouching Transformer with Limited Paired DataLe Jiang, Yan Huang, Lianxin Xie, Wen Xue et al.ACM MM 2024 · 1 citation
- Video Harmonization with Triplet Spatio-Temporal Variation PatternsZonghui Guo, Xinyu Han, Jie Zhang, Shiguang Shan et al.CVPR 2024
- DLFormer: Discrete Latent Transformer for Video InpaintingJingjing Ren, Qingqing Zheng, Yuanyuan Zhao, Xuemiao Xu et al.CVPR 2022 · 39 citations
- Blemish-aware and Progressive Face Retouching with Limited Paired DataLianxin Xie, Wen Xue, Zhen Xu, Si Wu et al.CVPR 2023
