Lune

CVPR2025顶会

From Faces to Voices: Learning Hierarchical Representations for High-quality Video-to-Speech

Ji-Hoon Kim, Jeongsoo Choi, Jaehun Kim, Chaeyoung Jung, Joon Son Chung

2025年份
5顶会引用

摘要

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as videoto-speech synthesis. A significant challenge in video-tospeech synthesis lies in the substantial modality gap between silent video and multi-faceted speech. In this paper, we propose a novel video-to-speech system that effectively bridges this modality gap, significantly enhancing the quality of synthesized speech. This is achieved by learning of hierarchical representations from video to speech. Specifically, we gradually transform silent video into acoustic feature spaces through three sequential stages -content, timbre, and prosody modeling. In each stage, we align visual factors -lip movements, face identity, and facial expressions -with corresponding acoustic counterparts to ensure the seamless transformation. Additionally, to generate realistic and coherent speech from the visual representations, we employ a flow matching model that estimates direct trajectories from a simple prior distribution to the target speech distribution. Extensive experiments demonstrate that our method achieves exceptional generation quality comparable to real utterances, outperforming existing methods by a significant margin.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext bf2a15ae-c31f-4bdf-bac6-ea908563ca7b

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper33

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖