DAVD-Net: Deep Audio-Aided Video Decompression of Talking Heads
Xi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben, Chengjie Tu
摘要
Close-up talking heads are among the most common and salient object in video contents, such as face-to-face conversations in social media, teleconferences, news broadcasting, talk shows, etc. Due to the high sensitivity of human visual system to faces, compression distortions in talking head videos are highly visible and annoying. To address this problem, we present a novel deep convolutional neural network method for very low bit rate video reconstruction of talking heads. The key innovation is a new DCNN architecture that can exploit the audio-video correlations to repair compression defects in the face region. We further improve reconstruction quality by embedding into our DCNN the encoder information of the video compression standards and introducing a constraining projection module in the network. Extensive experiments demonstrate that the proposed DCNN method outperforms the existing stateof-the-art methods on videos of talking heads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Video Coding using Learned Latent GAN CompressionMustafa Shukor, Bharath Bhushan Damodaran, Xu Yao, Pierre HellierACM MM 2022 · 被引用 6 次
- BADiff: Bandwidth Adaptive Diffusion ModelXi Zhang, Hanwei Zhu, Yan Zhong, Jiamang Wang 等NeurIPS 2025 · 被引用 1 次
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationShuai Shen, Wenliang Zhao, Zibin Meng, Wanhua Li 等CVPR 2023
- LipFormer: High-fidelity and Generalizable Talking Face Generation with A Pre-learned Facial CodebookJiayu Wang, Kang Zhao, Shiwei Zhang, Yingya Zhang 等CVPR 2023
它引用的顶会 Paper2
相关 Paper
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 被引用 9 次
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual PriorsSindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2022 · 被引用 1 次
- Talking Head from Speech Audio using a Pre-trained Image GeneratorMohammed M. Alghamdi, He Wang, Andrew J. Bulpitt, David C. HoggACM MM 2022 · 被引用 25 次
- Gemino: Practical and Robust Neural Compression for Video ConferencingVibhaalakshmi Sivaraman, Pantea Karimi, Vedantha Venkatapathy, Mehrdad Khani Shirkoohi 等NSDI 2024
- DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoZhimeng Zhang, Zhipeng Hu, Wenjin Deng, Changjie Fan 等AAAI 2023 · 被引用 106 次
