INRAS: Implicit Neural Representation for Audio Scenes
Kun Su, Mingfei Chen, Eli Shlizerman
Abstract
The spatial acoustic information of a scene, i.e., how sounds emitted from a particular location in the scene are perceived in another location, is key for immersive scene modeling. Robust representation of scene’s acoustics can be formulated through a continuous field formulation along with impulse responses varied by emitter-listener locations. The impulse responses are then used to render sounds perceived by the listener. While such representation is advantageous, parameterization of impulse responses for generic scenes presents itself as a challenge. Indeed, traditional pre-computation methods have only implemented parameterization at discrete probe points and require large storage, while other existing methods such as geometry-based sound simulations still suffer from inability to simulate all wave-based sound effects. In this work, we introduce a novel neural network for light-weight Implicit Neural Representation for Audio Scenes (INRAS), which can render a high fidelity time-domain impulse responses at any arbitrary emitter-listener positions by learning a continuous implicit function. INRAS disentangles scene’s geometry features with three modules to generate independent features for the emitter, the geometry of the scene, and the listener respectively. These lead to an efficient reuse of scene-dependent features and support effective multi-condition training for multiple scenes. Our experimental results show that INRAS outperforms existing approaches for representation and rendering of sounds for varying emitter-listener locations in all aspects, including the impulse response quality, inference speed, and storage requirements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6f6518f-9d65-49f8-b7b6-5a219016b251Cited by top-tier papers24
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar et al.NeurIPS 2023 · 77 citations
- NVRC: Neural Video Representation CompressionHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2024 · 44 citations
- Acoustic Volume Rendering for Neural Impulse Response FieldsZitong Lan, Chenhao Zheng, Zhiwei Zheng, Mingmin ZhaoNeurIPS 2024 · 35 citations
- AV-GS: Learning Material and Geometry Aware Priors for Novel View Acoustic SynthesisSwapnil Bhosale, Haosen Yang, Diptesh Kanojia, Jiankang Deng et al.NeurIPS 2024 · 22 citations
- AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud SplattingMingfei Chen, Eli ShlizermanNeurIPS 2024 · 14 citations
Builds on7
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Unsupervised Sound Separation Using Mixture Invariant TrainingScott Wisdom, Efthymios Tzinis, Hakan Erdogan, Ron J. Weiss et al.NeurIPS 2020 · 227 citations
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum et al.NeurIPS 2022 · 153 citations
- Audeo: Audio Generation for a Silent Performance VideoKun Su, Xiulong Liu, Eli ShlizermanNeurIPS 2020 · 78 citations
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn et al.ICLR 2021 · 73 citations
Related papers
- Neural Experts: Mixture of Experts for Implicit Neural RepresentationsYizhak Ben-Shabat, Chamin Hewa Koneputugodage, Sameera Ramasinghe, Stephen GouldNeurIPS 2024 · 14 citations
- NeRAF: 3D Scene Infused Neural Radiance and Acoustic FieldsAmandine Brunetto, Sascha Hornauer, Fabien MoutardeICLR 2025
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum et al.NeurIPS 2021 · 426 citations
- MESH2IR: Neural Acoustic Impulse Response Generator for Complex 3D ScenesAnton Ratnarajah, Zhenyu Tang, Rohith Aralikatti, Dinesh ManochaACM MM 2022 · 35 citations
- VI^3NR: Variance Informed Initialization for Implicit Neural RepresentationsChamin Hewa Koneputugodage, Yizhak Ben-Shabat, Sameera Ramasinghe, Stephen GouldCVPR 2025
