Vision-Infused Deep Audio Inpainting
Hang Zhou, Ziwei Liu, Xudong Xu, Ping Luo, Xiaogang Wang
Abstract
Multi-modality perception is essential to develop interactive intelligence. In this work, we consider a new task of visual information-infused audio inpainting, i.e. synthesizing missing audio segments that correspond to their accompanying videos. We identify two key aspects for a successful inpainter: (1) It is desirable to operate on spectrograms instead of raw audios. Recent advances in deep semantic image inpainting could be leveraged to go beyond the limitations of traditional audio inpainting. (2) To synthesize visually indicated audio, a visualaudio joint feature space needs to be learned with synchronization of audio and video. To facilitate a largescale study, we collect a new multi-modality instrumentplaying dataset called MUSIC-Extra-Solo (MUSICES) by enriching MUSIC dataset [54] . Extensive experiments demonstrate that our framework is capable of inpainting realistic and varying audio segments with or without visual contexts. More importantly, our synthesized audio segments are coherent with their video counterparts, showing the effectiveness of our proposed Vision-Infused Audio Inpainter (VIAI). Code, models, dataset and video results are available at https://hangz-nju-cuhk. github.io/projects/AudioInpainting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3bfb8a32-b880-4419-aff1-ca7c8a8a4341Cited by top-tier papers26
- AUDIT: Audio Editing by Following Instructions with Latent Diffusion ModelsYuancheng Wang, Zeqian Ju, Xu Tan, Lei He et al.NeurIPS 2023 · 120 citations
- Expressive Talking Head Generation with Granular Audio-Visual ControlBorong Liang, Yan Pan, Zhizhi Guo, Hang Zhou et al.CVPR 2022 · 114 citations
- Exploiting Explanations for Model Inversion AttacksXuejun Zhao, Wencan Zhang, Xiaokui Xiao, Brian Y. LimICCV 2021 · 113 citations
- Learning to Answer Questions in Dynamic Audio-Visual ScenariosGuangyao Li, Yake Wei, Yapeng Tian, Chenliang Xu et al.CVPR 2022 · 101 citations
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 80 citations
Builds on2
Related papers
- MPJudge: Towards Perceptual Assessment of Music-Induced PaintingsShiqi Jiang, Tianyi Liang, Huayuan Ye, Changbo Wang et al.AAAI 2026
- Distilling Audio-Visual Knowledge by Compositional Contrastive LearningYanbei Chen, Yongqin Xian, A. Sophia Koepke, Ying Shan et al.CVPR 2021
- Audio-Visual Instance SegmentationRuohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu et al.CVPR 2025
- An Adaptive Iterative Inpainting Method with More Information ExplorationShengjie Chen, Zhenhua Guo, Bo YuanACM MM 2021 · 2 citations
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 4 citations
