DAVD-Net: Deep Audio-Aided Video Decompression of Talking Heads
Xi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben, Chengjie Tu
Abstract
Close-up talking heads are among the most common and salient object in video contents, such as face-to-face conversations in social media, teleconferences, news broadcasting, talk shows, etc. Due to the high sensitivity of human visual system to faces, compression distortions in talking head videos are highly visible and annoying. To address this problem, we present a novel deep convolutional neural network method for very low bit rate video reconstruction of talking heads. The key innovation is a new DCNN architecture that can exploit the audio-video correlations to repair compression defects in the face region. We further improve reconstruction quality by embedding into our DCNN the encoder information of the video compression standards and introducing a constraining projection module in the network. Extensive experiments demonstrate that the proposed DCNN method outperforms the existing stateof-the-art methods on videos of talking heads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Video Coding using Learned Latent GAN CompressionMustafa Shukor, Bharath Bhushan Damodaran, Xu Yao, Pierre HellierACM MM 2022 · 6 citations
- BADiff: Bandwidth Adaptive Diffusion ModelXi Zhang, Hanwei Zhu, Yan Zhong, Jiamang Wang et al.NeurIPS 2025 · 1 citation
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationShuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li et al.CVPR 2023
- LipFormer: High-fidelity and Generalizable Talking Face Generation with A Pre-learned Facial CodebookJiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang et al.CVPR 2023
Builds on2
Related papers
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 9 citations
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual PriorsSindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2022 · 1 citation
- Talking Head from Speech Audio using a Pre-trained Image GeneratorMohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. HoggACM MM 2022 · 25 citations
- Gemino: Practical and Robust Neural Compression for Video ConferencingVibhaalakshmi Sivaraman, Pantea Karimi, Vedantha Venkatapathy, Mehrdad Khani Shirkoohi et al.NSDI 2024
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan et al.AAAI 2023 · 106 citations
