Extreme-scale Talking-Face Video Upsampling with Audio-Visual Priors
Sindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar
摘要
In this paper, we explore an interesting question of what can be obtained from an 8x8 pixel video sequence. Surprisingly, it turns out to be quite a lot. We show that when we process this 8x8 video with the right set of audio and image priors, we can obtain a full-length, 256x256 video. We achieve this 32x scaling of an extremely low-resolution input using our novel audio-visual upsampling network. The audio prior helps to recover the elemental facial details and precise lip shapes and a single high-resolution target identity image prior provides us with rich appearance details. Our approach is an end-to-end multi-stage framework. The first stage produces a coarse intermediate output video that can be then used to animate single target identity image and generate realistic, accurate and high-quality outputs. Our approach is simple and performs exceedingly well (an 8x improvement in FID score) compared to previous super-resolution methods. We also extend our model to talking-face video compression, and show that we obtain a 3.5x improvement in terms of bits/pixel over the previous state-of-the-art. The results from our network are thoroughly analyzed through extensive ablation experiments (in the paper and supplementary material). We also provide the demo video along with code and models on our http://cvit.iiit.ac.in/research/projects/cvit-projects/talking-face-video-upsampling.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 被引用 869 次
- Learning temporal coherence via self-supervision for GAN-based video generationMengyu Chu, You Xie, Jonas Mayer, Laura Leal-Taixé 等SIGGRAPH 2020 · 被引用 198 次
- FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningChenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng 等ICCV 2021 · 被引用 149 次
- Joint Implicit Image Function for Guided Depth Super-ResolutionJiaxiang Tang, Xiaokang Chen, Gang ZengACM MM 2021 · 被引用 78 次
- Efficient Video Compression via Content-Adaptive Super-ResolutionMehrdad Khani Shirkoohi, Vibhaalakshmi Sivaraman, Mohammad AlizadehICCV 2021 · 被引用 68 次
相关 Paper
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben 等CVPR 2020
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 被引用 9 次
- Learning to Have an Ear for Face Super-ResolutionGivi Meishvili, Simon Jenni, Paolo FavaroCVPR 2020
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine PerceptionYudong Mao, Peilin Chen, Shurun Wang, Shiqi Wang 等ACM MM 2023 · 被引用 5 次
