Lune

ICCV2019Top-tier venue

Vision-Infused Deep Audio Inpainting

Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, Xiaogang Wang

2019Year
92Citations
26Top-tier citations

Abstract

Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, i.e. synthesizing missing audio segments that correspond to their accompanying videos. We identify two key aspects for a successful inpainter: (1) It is desirable to operate on spectrograms instead of raw audios. Recent advances in deep semantic image inpainting could be leveraged to go beyond the limitations of traditional audio inpainting. (2) To synthesize visually indicated audio, a visualaudio joint feature space needs to be learned with synchronization of audio and video. To facilitate a largescale study, we collect a new multi-modality instrumentplaying dataset called MUSIC-Extra-Solo (MUSICES) by enriching MUSIC dataset [54] . Extensive experiments demonstrate that our framework is capable of inpainting realistic and varying audio segments with or without visual contexts. More importantly, our synthesized audio segments are coherent with their video counterparts, showing the effectiveness of our proposed Vision-Infused Audio Inpainter (VIAI). Code, models, dataset and video results are available at https://hangz-nju-cuhk. github.io/projects/AudioInpainting .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3bfb8a32-b880-4419-aff1-ca7c8a8a4341

Cited by top-tier papers26

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines