Show and Polish: Reference-Guided Identity Preservation in Face Video Restoration
Wenkang Han, Wang Lin, Yiyun Zhou, Qi Liu, Shulei Wang, Chang Yao, Jingyuan Chen
Abstract
Face Video Restoration (FVR) aims to reconstruct high-quality face videos from degraded input. Traditional methods struggle to preserve fine-grained, identity-specific features when degradation is severe, often producing average-looking faces that lack individual characteristics. To address these challenges, we introduce IP-FVR, a novel method that leverages a high-quality reference face image as a visual prompt to provide identity conditioning during the denoising process. IP-FVR incorporates semantically rich identity information from the reference image using decoupled cross-attention mechanisms, ensuring detailed and identity consistent results. For intra-clip identity drift (within 24 frames), we introduce an identity-preserving feedback learning method that combines cosine similarity-based reward signals with suffix-weighted temporal aggregation. This approach effectively minimizes drift within sequences of frames. For inter-clip identity drift, we develop an exponential blending strategy that aligns identities across clips by iteratively blending frames from previous clips during the denoising process. This method ensures consistent identity representation across different clips. Additionally, we enhance the restoration process with a multi-stream negative prompt, guiding the model's attention to relevant facial attributes and minimizing the generation of low-quality or incorrect features. Extensive experiments on both synthetic and real-world datasets demonstrate that IP-FVR outperforms existing methods in both quality and identity preservation, showcasing its substantial potential for practical applications in face video restoration. Our code and datasets are available at https://ip-fvr.github.io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency RewardsShulei Wang, Longhui Wei, Xin He, Jianbo Ouyang et al.CVPR 2026 · 7 citations
- TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather RemovalHanting Wang, Shengpeng Ji, Shulei Wang, Hai Huang et al.ACM MM 2025
- DynaMem: Consistent Long Video Generation via Hierarchical Memory and Motion PriorsJingyu Lin, Xinyi Shang, Peng Sun, Cunjian Chen et al.ICML 2026
Builds on41
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar et al.ICCV 2021 · 1,325 citations
Related papers
- Blind Face Video Restoration with Temporal Consistent Generative Prior and Degradation-Aware PromptJingfan Tan, Hyunhee Park, Ying Zhang, Tao Wang et al.ACM MM 2024 · 10 citations
- Dynamic Content Prediction with Motion-aware Priors for Blind Face Video RestorationLianxin Xie, Bingbing Zheng, Si Wu, Hau-San WongCVPR 2025
- Face Video Deblurring Using 3D Facial PriorsWenqi Ren, Jiaolong Yang, Senyou Deng, David P. Wipf et al.ICCV 2019 · 52 citations
- Identity-Preserving Image-to-Video Generation via Reward-Guided OptimizationLiao Shen, Wentao Jiang, Yiran Zhu, Jiahe Li et al.CVPR 2026 · 8 citations
- FaceMe: Robust Blind Face Restoration with Personal IdentificationSiyu Liu, Zheng-Peng Duan, Jia Ouyang, Jiayi Fu et al.AAAI 2025 · 18 citations
