AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis
Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar, Chenliang Xu
摘要
Can machines recording an audio-visual scene produce realistic, matching audiovisual experiences at novel positions and novel view directions? We answer it by studying a new task-real-world audio-visual scene synthesis-and a first-of-itskind NeRF-based approach for multimodal learning. Concretely, given a video recording of an audio-visual scene, the task is to synthesize new videos with spatial audios along arbitrary novel camera trajectories in that scene. We propose an acoustic-aware audio generation module that integrates prior knowledge of audio propagation into NeRF, in which we implicitly associate audio generation with the 3D geometry and material properties of a visual environment. Furthermore, we present a coordinate transformation module that expresses a view direction relative to the sound source, enabling the model to learn sound source-centric acoustic fields. To facilitate the study of this new task, we collect a high-quality Real-World Audio-Visual Scene (RWAVS) dataset. We demonstrate the advantages of our method on this real-world dataset and the simulation-based SoundSpaces dataset. We recommend that readers visit our project page for convincing comparisons: https://liangsusan-git.github.io/project/avnerf/ . Introduction We study a new task, real-world audio-visual scene synthesis, to generate target videos and audios along novel camera trajectories from source audio-visual recordings of known trajectories. By learning from real-world source videos with binaural audio, we aim to generate target video frames and spatial audios that exhibit consistency with the given camera trajectory visually and acoustically. This consistency ensures perceptual realism and immersion, enriching the overall user experience. As far as we know, attempts in the audio-visual learning literature [1-11] have yet to succeed in solving this challenging task thus far. Although there are similar works [12] [13] [14] [15] , these methods have constraints that limit their ability to solve this new task. Luo et al. [12] propose neural acoustic fields to model sound propagation in a room. Su et al. [13] introduce representing audio scenes by disentangling the scene's geometry features. These methods are tailored for estimating room impulse response signals in a simulation environment that are difficult to obtain in a real-world scene. Concurrent to our work, ViGAS proposed by Chen et al. [15] learns to synthesize new sounds by inferring the audio-visual cues. However, ViGAS is limited to a few viewpoints for audio generation. We introduce AV-NeRF, a novel NeRF-based method of synthesizing real-world audio-visual scenes. AV-NeRF enables the generation of videos and spatial audios, following arbitrary camera trajectories. It utilizes source videos and camera poses as references. AV-NeRF consists of two branches: A-NeRF, which learns the acoustic fields of an environment, and V-NeRF, which models color and density fields. We represent a static audio field as a continuous function using A-NeRF, which takes the listener's position and head direction as input. A-NeRF effectively models the energy decay of sound as the sound travels from the source to the listener by correlating the listener's position with the 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Acoustic Volume Rendering for Neural Impulse Response FieldsZitong Lan, Chenhao Zheng, Zhiwei Zheng, Mingmin ZhaoNeurIPS 2024 · 被引用 35 次
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 被引用 22 次
- AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic SynthesisSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng 等NeurIPS 2024 · 被引用 22 次
- Sound Localization from Motion: Jointly Learning Sound Direction and Camera RotationZiyang Chen, Shengyi Qian, Andrew OwensICCV 2023 · 被引用 21 次
- AV-RIR: Audio-Visual Room Impulse Response EstimationAnton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya 等CVPR 2024 · 被引用 15 次
它引用的顶会 Paper23
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun 等NeurIPS 2020 · 被引用 1,010 次
- Nerfstudio: A Modular Framework for Neural Radiance Field DevelopmentMatthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li 等SIGGRAPH 2023 · 被引用 592 次
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 被引用 522 次
相关 Paper
- NeRAF: 3D Scene Infused Neural Radiance and Acoustic FieldsAmandine Brunetto, Sascha Hornauer, Fabien MoutardeICLR 2025
- Novel-View Acoustic SynthesisChangan Chen, Alexander Richard, Roman Shapovalov, Vamsi Krishna Ithapu 等CVPR 2023
- Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and BenchmarkZiyang Chen, Israel D. Gebru, Christian Richardt, Anurag Kumar 等CVPR 2024
- Multimodal Neural Acoustic Fields for Immersive Mixed RealityGuaneen Tong, Johnathan Chi-Ho Leung, Xi Peng, Haosheng Shi 等IEEE VR 2025 · 被引用 4 次
- -AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?Susan Liang, Chao Huang, Yunlong Tang, Zeliang Zhang 等ICCV 2025 · 被引用 4 次
