Lune

CVPR2025Top-tier venue

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung, Joon Son Chung

2025Year
5Top-tier citations

Abstract

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as videoto-speech synthesis. A significant challenge in video-tospeech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we propose a novel video-to-speech system that effectively bridges this modality gap, significantly enhancing the quality of synthesized speech. This is achieved by learning of hierarchical representations from video to speech. Specifically, we gradually transform silent video into acoustic feature spaces through three sequential stages -content, timbre, and prosody modeling. In each stage, we align visual factors -lip movements, face identity, and facial expressions -with corresponding acoustic counterparts to ensure the seamless transformation. Additionally, to generate realistic and coherent speech from the visual representations, we employ a flow matching model that estimates direct trajectories from a simple prior distribution to the target speech distribution. Extensive experiments demonstrate that our method achieves exceptional generation quality comparable to real utterances, outperforming existing methods by a significant margin.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bf2a15ae-c31f-4bdf-bac6-ea908563ca7b

Cited by top-tier papers5

Ask how each one uses it

Builds on33

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines