Flow-Based Unconstrained Lip to Speech Generation
Jinzheng He, Zhou Zhao, Yi Ren, Jinglin Liu, Baoxing Huai, Nicholas Jing Yuan
摘要
Unconstrained lip-to-speech aims to generate corresponding speeches based on silent facial videos with no restriction to head pose or vocabulary. It is desirable to generate intelligible and natural speech with a fast speed in unconstrained settings. Currently, to handle the more complicated scenarios, most existing methods adopt the autoregressive architecture, which is optimized with the MSE loss. Although these methods have achieved promising performance, they are prone to bring issues including high inference latency and mel-spectrogram over-smoothness. To tackle these problems, we propose a novel flow-based non-autoregressive lip-to-speech model (GlowLTS) to break autoregressive constraints and achieve faster inference. Concretely, we adopt a flow-based decoder which is optimized by maximizing the likelihood of the training data and is capable of more natural and fast speech generation. Moreover, we devise a condition module to improve the intelligibility of generated speech. We demonstrate the superiority of our proposed method through objective and subjective evaluation on Lip2Wav-Chemistry-Lectures and Lip2Wav-Chess-Analysis datasets. Our demo video can be found at https://glowlts.github.io/.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Let There Be Sound: Reconstructing High Quality Speech from Silent VideosJi-Hoon Kim, Jaehun Kim, Joon Son ChungAAAI 2024 · 被引用 14 次
- CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-TrainingZhenhui Ye, Rongjie Huang, Yi Ren, Ziyue Jiang 等ACL 2023 · 被引用 13 次
- On the Robustness of Normalizing Flows for Inverse Problems in ImagingSeongmin Hong, Inbum Park, Se Young ChunICCV 2023 · 被引用 9 次
- FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech SynthesisYongqi Wang, Zhou ZhaoACM MM 2022 · 被引用 9 次
- Neural Diffeomorphic Non-uniform B-spline FlowsSeongmin Hong, Se Young ChunAAAI 2023 · 被引用 3 次
它引用的顶会 Paper3
- Glow-TTS: A Generative Flow for Text-to-Speech via Monotonic Alignment SearchJaehyeon Kim, Sungwon Kim, Jungil Kong, Sungroh YoonNeurIPS 2020 · 被引用 663 次
- Why Normalizing Flows Fail to Detect Out-of-Distribution DataPolina Kirichenko, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2020 · 被引用 370 次
- C-Flow: Conditional Generative Flow Models for Images and 3D Point CloudsAlbert Pumarola, Stefan Popov, Francesc Moreno-Noguer, Vittorio FerrariCVPR 2020
相关 Paper
- FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANsAndreas Zinonos, Michał Stypułkowski, Antoni Bigata Casademunt, Stavros Petridis 等CVPR 2026
- FastLR: Non-Autoregressive Lipreading Model with Integrate-and-FireJinglin Liu, Yi Ren, Zhou Zhao, Chen Zhang 等ACM MM 2020 · 被引用 13 次
- Speech2Lip: High-fidelity Speech to Lip Generation by Learning from a Short VideoXiuzhe Wu, Pengfei Hu, Yang Wu, Xiaoyang Lyu 等ICCV 2023 · 被引用 18 次
- Lip-to-Speech Synthesis for Arbitrary Speakers in the WildSindhu B. Hegde, K. R. Prajwal, Rudrabha Mukhopadhyay, Vinay P. Namboodiri 等ACM MM 2022 · 被引用 15 次
- Bidirectional Variational Inference for Non-Autoregressive Text-to-SpeechYoonhyung Lee, Joongbo Shin, Kyomin JungICLR 2021 · 被引用 42 次
