AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene Synthesis
Susan Liang, Chao Huang, Yapeng Tian, Anurag Kumar, Chenliang Xu
Abstract
Can machines recording an audio-visual scene produce realistic, matching audiovisual experiences at novel positions and novel view directions? We answer it by studying a new task-real-world audio-visual scene synthesis-and a first-of-itskind NeRF-based approach for multimodal learning. Concretely, given a video recording of an audio-visual scene, the task is to synthesize new videos with spatial audios along arbitrary novel camera trajectories in that scene. We propose an acoustic-aware audio generation module that integrates prior knowledge of audio propagation into NeRF, in which we implicitly associate audio generation with the 3D geometry and material properties of a visual environment. Furthermore, we present a coordinate transformation module that expresses a view direction relative to the sound source, enabling the model to learn sound source-centric acoustic fields. To facilitate the study of this new task, we collect a high-quality Real-World Audio-Visual Scene (RWAVS) dataset. We demonstrate the advantages of our method on this real-world dataset and the simulation-based SoundSpaces dataset. We recommend that readers visit our project page for convincing comparisons: https://liangsusan-git.github.io/project/avnerf/ . Introduction We study a new task, real-world audio-visual scene synthesis, to generate target videos and audios along novel camera trajectories from source audio-visual recordings of known trajectories. By learning from real-world source videos with binaural audio, we aim to generate target video frames and spatial audios that exhibit consistency with the given camera trajectory visually and acoustically. This consistency ensures perceptual realism and immersion, enriching the overall user experience. As far as we know, attempts in the audio-visual learning literature [1-11] have yet to succeed in solving this challenging task thus far. Although there are similar works [12] [13] [14] [15] , these methods have constraints that limit their ability to solve this new task. Luo et al. [12] propose neural acoustic fields to model sound propagation in a room. Su et al. [13] introduce representing audio scenes by disentangling the scene's geometry features. These methods are tailored for estimating room impulse response signals in a simulation environment that are difficult to obtain in a real-world scene. Concurrent to our work, ViGAS proposed by Chen et al. [15] learns to synthesize new sounds by inferring the audio-visual cues. However, ViGAS is limited to a few viewpoints for audio generation. We introduce AV-NeRF, a novel NeRF-based method of synthesizing real-world audio-visual scenes. AV-NeRF enables the generation of videos and spatial audios, following arbitrary camera trajectories. It utilizes source videos and camera poses as references. AV-NeRF consists of two branches: A-NeRF, which learns the acoustic fields of an environment, and V-NeRF, which models color and density fields. We represent a static audio field as a continuous function using A-NeRF, which takes the listener's position and head direction as input. A-NeRF effectively models the energy decay of sound as the sound travels from the source to the listener by correlating the listener's position with the 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1c8e8bc2-c0d2-4f07-b9a8-e8932bb5f3adCited by top-tier papers29
- Acoustic Volume Rendering for Neural Impulse Response FieldsZitong Lan, Chenhao Zheng, Zhiwei Zheng, Mingmin ZhaoNeurIPS 2024 · 35 citations
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 22 citations
- AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic SynthesisSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng et al.NeurIPS 2024 · 22 citations
- Sound Localization from Motion: Jointly Learning Sound Direction and Camera RotationZiyang Chen, Shengyi Qian, Andrew OwensICCV 2023 · 21 citations
- AV-RIR: Audio-Visual Room Impulse Response EstimationAnton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya et al.CVPR 2024 · 15 citations
Builds on23
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra et al.ICCV 2019 · 1,863 citations
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun et al.NeurIPS 2020 · 1,010 citations
- Nerfstudio: A Modular Framework for Neural Radiance Field DevelopmentMatthew Tancik, Ethan Weber, Evonne Ng, Ruilong Li et al.SIGGRAPH 2023 · 592 citations
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 522 citations
Related papers
- NeRAF: 3D Scene Infused Neural Radiance and Acoustic FieldsAmandine Brunetto, Sascha Hornauer, Fabien MoutardeICLR 2025
- Novel-View Acoustic SynthesisChangan Chen, Alexander Richard, Roman Shapovalov, Vamsi Krishna Ithapu et al.CVPR 2023
- Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and BenchmarkZiyang Chen, Israel D. Gebru, Christian Richardt, Anurag Kumar et al.CVPR 2024
- Multimodal Neural Acoustic Fields for Immersive Mixed RealityGuaneen Tong, Johnathan Chi-Ho Leung, Xi Peng, Haosheng Shi et al.IEEE VR 2025 · 4 citations
- -AVAS: Can Physics-Integrated Audio-Visual Modeling Boost Neural Acoustic Synthesis?Susan Liang, Chao Huang, Yunlong Tang, Zeliang Zhang et al.ICCV 2025 · 4 citations
