Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
Yutong Wang, Jiajie Teng, Jiajiong Cao, Yuming Li, Chenguang Ma, Hongteng Xu, Dixin Luo
Abstract
As a very common type of video, face videos often appear in movies, talk shows, live broadcasts, and other scenes. Real-world online videos are often plagued by degradations such as blurring and quantization noise, due to the high compression ratio caused by high communication costs and limited transmission bandwidth. These degradations have a particularly serious impact on face videos because the human visual system is highly sensitive to facial details. Despite the significant advancement in video face enhancement, current methods still suffer from i) long processing time and ii) inconsistent spatial-temporal visual effects (e.g., flickering). This study proposes a novel and efficient blind video face enhancement method to overcome the above two challenges, restoring high-quality videos from their compressed low-quality versions with an effective deflickering mechanism. In particular, the proposed method develops upon a 3D-VQGAN backbone associated with spatial-temporal codebooks recording high-quality portrait features and residual-based temporal information. We develop a two-stage learning framework for the model. In Stage I, we learn the model with a regularizer mitigating the codebook collapse problem. In Stage II, we learn two transformers to look up code from the codebooks and further update the encoder of low-quality videos. Experiments conducted on the VFHQ-Test dataset demonstrate that our method surpasses the current state-of-the-art blind face video restoration and de-flickering methods on both efficiency and effectiveness. Code is available at https: //github.com/Dixin-Lab/BFVR-STC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on22
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- vq-wav2vec: Self-Supervised Learning of Discrete Speech RepresentationsAlexei Baevski, Steffen Schneider, Michael AuliICLR 2020 · 730 citations
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Towards Robust Blind Face Restoration with Codebook Lookup TransformerShangchen Zhou, Kelvin C. K. Chan, Chongyi Li, Chen Change LoyNeurIPS 2022 · 431 citations
- MoVQ: Modulating Quantized Vectors for High-Fidelity Image GenerationChuanxia Zheng, Tung-Long Vuong, Jianfei Cai, Dinh PhungNeurIPS 2022 · 156 citations
Related papers
- Dirichlet-Constrained Variational Codebook Learning for Temporally Coherent Video Face RestorationBaoyou Chen, Ce Liu, Weihao Yuan, Zilong Dong et al.ICCV 2025 · 3 citations
- LD-BFR: Vector-Quantization-Based Face Restoration Model with Latent Diffusion EnhancementYuzhen Du, Teng Hu, Ran Yi, Lizhuang MaACM MM 2024 · 3 citations
- Exploring Correlations in Degraded Spatial Identity Features for Blind Face RestorationQian Ning, Fangfang Wu, Weisheng Dong, Xin Li et al.ACM MM 2023 · 1 citation
- Discrete Prior-Based Temporal-Coherent Content Prediction for Blind Face Video RestorationLianxin Xie, Bingbing Zheng, Wen Xue, Yunfei Zhang et al.AAAI 2025
- Audio-Assisted Face Video Restoration with Temporal and Identity Complementary LearningYuqin Cao, Yixuan Gao, Wei Sun, Xiaohong Liu et al.AAAI 2026
