Lune

ICCV2019顶会

Vision-Infused Deep Audio Inpainting

Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, Xiaogang Wang

2019年份
92被引次数
26顶会引用

摘要

Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, i.e. synthesizing missing audio segments that correspond to their accompanying videos. We identify two key aspects for a successful inpainter: (1) It is desirable to operate on spectrograms instead of raw audios. Recent advances in deep semantic image inpainting could be leveraged to go beyond the limitations of traditional audio inpainting. (2) To synthesize visually indicated audio, a visualaudio joint feature space needs to be learned with synchronization of audio and video. To facilitate a largescale study, we collect a new multi-modality instrumentplaying dataset called MUSIC-Extra-Solo (MUSICES) by enriching MUSIC dataset [54] . Extensive experiments demonstrate that our framework is capable of inpainting realistic and varying audio segments with or without visual contexts. More importantly, our synthesized audio segments are coherent with their video counterparts, showing the effectiveness of our proposed Vision-Infused Audio Inpainter (VIAI). Code, models, dataset and video results are available at https://hangz-nju-cuhk. github.io/projects/AudioInpainting .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 3bfb8a32-b880-4419-aff1-ca7c8a8a4341

引用它的顶会 Paper26

问问它们各自怎么用它

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖