Extreme-scale Talking-Face Video Upsampling with Audio-Visual Priors
Sindhu B. Hegde, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. Jawahar
Abstract
In this paper, we explore an interesting question of what can be obtained from an 8x8 pixel video sequence. Surprisingly, it turns out to be quite a lot. We show that when we process this 8x8 video with the right set of audio and image priors, we can obtain a full-length, 256x256 video. We achieve this 32x scaling of an extremely low-resolution input using our novel audio-visual upsampling network. The audio prior helps to recover the elemental facial details and precise lip shapes and a single high-resolution target identity image prior provides us with rich appearance details. Our approach is an end-to-end multi-stage framework. The first stage produces a coarse intermediate output video that can be then used to animate single target identity image and generate realistic, accurate and high-quality outputs. Our approach is simple and performs exceedingly well (an 8x improvement in FID score) compared to previous super-resolution methods. We also extend our model to talking-face video compression, and show that we obtain a 3.5x improvement in terms of bits/pixel over the previous state-of-the-art. The results from our network are thoroughly analyzed through extensive ablation experiments (in the paper and supplementary material). We also provide the demo video along with code and models on our http://cvit.iiit.ac.in/research/projects/cvit-projects/talking-face-video-upsampling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a64e5ec8-d2de-48e9-8ab6-01b6792eb9aeBuilds on11
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In the WildK. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri, C. V. JawaharACM MM 2020 · 869 citations
- Learning temporal coherence via self-supervision for GAN-based video generationMengyu Chu, You Xie, Jonas Mayer, Laura Leal-Taixé et al.SIGGRAPH 2020 · 198 citations
- FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningChenxu Zhang, Yifan Zhao, Yifei Huang, Ming Zeng et al.ICCV 2021 · 149 citations
- Joint Implicit Image Function for Guided Depth Super-ResolutionJiaxiang Tang, Xiaokang Chen, Gang ZengACM MM 2021 · 78 citations
- Efficient Video Compression via Content-Adaptive Super-ResolutionMehrdad Khani Shirkoohi, Vibhaalakshmi Sivaraman, Mohammad AlizadehICCV 2021 · 68 citations
Related papers
- DAVD-Net: Deep Audio-Aided Video Decompression of Talking HeadsXi Zhang, Xiaolin Wu, Xinliang Zhai, Xianye Ben et al.CVPR 2020
- Deep Multi-modality Soft-decoding of Very Low Bit-rate Face VideosYanhui Guo, Xi Zhang, Xiaolin WuACM MM 2020 · 9 citations
- Learning to Have an Ear for Face Super-ResolutionGivi Meishvili, Simon Jenni, Paolo FavaroCVPR 2020
- One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingTing-Chun Wang, Arun Mallya, Ming-Yu LiuCVPR 2021
- Peering into The Sketch: Ultra-Low Bitrate Face Compression for Joint Human and Machine PerceptionYudong Mao, Peilin Chen, Shurun Wang, Shiqi Wang et al.ACM MM 2023 · 5 citations
