Deep Multi-modality Soft-decoding of Very Low Bit-rate Face Videos
Yanhui Guo, Xi Zhang, Xiaolin Wu
摘要
We propose a novel deep multi-modality neural network for restoring very low bit rate videos of talking heads. Such video contents are very common in social media, teleconferencing, distance education, tele-medicine, etc., and often need to be transmitted with limited bandwidth. The proposed CNN method exploits the correlations among three modalities, video, audio and emotion state of the speaker, to remove the video compression artifacts caused by spatial down sampling and quantization. The deep learning approach turns out to be ideally suited for the video restoration task, as the complex non-linear cross-modality correlations are very difficult to model analytically and explicitly. The new method is a video post processor that can significantly boost the perceptual quality of aggressively compressed talking head videos, while being fully compatible with all existing video compression standards.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Functional Neural Networks for Parametric Image Restoration ProblemsFangzhou Luo, Xiaolin Wu, Yanhui GuoNeurIPS 2021 · 被引用 21 次
- Audio-Assisted Face Video Restoration with Temporal and Identity Complementary LearningYuqin Cao, Yixuan Gao, Wei Sun, Xiaohong Liu 等AAAI 2026
- Learning Degradation-Independent Representations for Camera ISP PipelinesYanhui Guo, Fangzhou Luo, Xiaolin WuCVPR 2024
它引用的顶会 Paper3
- Non-Local ConvLSTM for Video Compression Artifact ReductionYi Xu, Longwen Gao, Kai Tian, Shuigeng Zhou 等ICCV 2019 · 被引用 70 次
- TDAN: Temporally-Deformable Alignment Network for Video Super-ResolutionYapeng Tian, Yulun Zhang, Yun Fu, Chenliang XuCVPR 2020
- MMTM: Multimodal Transfer Module for CNN FusionHamid Reza Vaezi Joze, Amirreza Shaban, Michael L. Iuzzolino, Kazuhito KoishidaCVPR 2020
相关 Paper
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben 等CVPR 2020
- Increasing Video Perceptual Quality with GANs and Semantic CodingLeonardo Galteri, Marco Bertini, Lorenzo Seidenari, Tiberio Uricchio 等ACM MM 2020 · 被引用 15 次
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual PriorsSindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2022 · 被引用 1 次
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Frequency-Biased Synergistic Design for Image Compression and CompensationJiaming Liu, Qi Zheng, Zihao Liu, Yilian Zhong 等CVPR 2025
