Deep Multi-modality Soft-decoding of Very Low Bit-rate Face Videos
Yanhui Guo, Xi Zhang, Xiaolin Wu
Abstract
We propose a novel deep multi-modality neural network for restoring very low bit rate videos of talking heads. Such video contents are very common in social media, teleconferencing, distance education, tele-medicine, etc., and often need to be transmitted with limited bandwidth. The proposed CNN method exploits the correlations among three modalities, video, audio and emotion state of the speaker, to remove the video compression artifacts caused by spatial down sampling and quantization. The deep learning approach turns out to be ideally suited for the video restoration task, as the complex non-linear cross-modality correlations are very difficult to model analytically and explicitly. The new method is a video post processor that can significantly boost the perceptual quality of aggressively compressed talking head videos, while being fully compatible with all existing video compression standards.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Functional Neural Networks for Parametric Image Restoration ProblemsFangzhou Luo, Xiaolin Wu, Yanhui GuoNeurIPS 2021 · 21 citations
- Audio-Assisted Face Video Restoration with Temporal and Identity Complementary LearningYuqin Cao, Yixuan Gao, Wei Sun, Xiaohong Liu et al.AAAI 2026
- Learning Degradation-Independent Representations for Camera ISP PipelinesYanhui Guo, Fangzhou Luo, Xiaolin WuCVPR 2024
Builds on3
- Non-Local ConvLSTM for Video Compression Artifact ReductionYi Xu, Longwen Gao, Kai Tian, Shuigeng Zhou et al.ICCV 2019 · 70 citations
- TDAN: Temporally-Deformable Alignment Network for Video Super-ResolutionYapeng Tian, Yulun Zhang, Yun Fu, Chenliang XuCVPR 2020
- MMTM: Multimodal Transfer Module for CNN FusionHamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito KoishidaCVPR 2020
Related papers
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben et al.CVPR 2020
- Increasing Video Perceptual Quality with GANs and Semantic CodingLeonardo Galteri, Marco Bertini, Lorenzo Seidenari, Tiberio Uricchio et al.ACM MM 2020 · 15 citations
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual PriorsSindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2022 · 1 citation
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Frequency-Biased Synergistic Design for Image Compression and CompensationJiaming Liu, Qi Zheng, Zihao Liu, Yilian Zhong et al.CVPR 2025
