Hearing Anywhere in Any Environment
Xiulong Liu, Anurag Kumar, Paul Calamia, Sebastià Vicenc Amengual Garí, Calvin Murdock, Ishwarya Ananthabhotla, Philip W. Robinson, Eli Shlizerman, Vamsi Krishna Ithapu, Ruohan Gao
摘要
In mixed reality applications, a realistic acoustic experience in spatial environments is as crucial as the visual experience for achieving true immersion. Despite recent advances in neural approaches for Room Impulse Response (RIR) estimation, most existing methods are limited to the single environment on which they are trained, lacking the ability to generalize to new rooms with different geometries and surface materials. We aim to develop a unified model capable of reconstructing the spatial acoustic experience of any environment with minimum additional measurements. To this end, we present XRIR, a framework for cross-room RIR prediction. The core of our generalizable approach lies in combining a geometric feature extractor, which captures spatial context from panorama depth images, with a RIR encoder that extracts detailed acoustic features from only a few reference RIR samples. To evaluate our method, we introduce ACOUSTICROOMS, a new dataset featuring highfidelity simulation of over 300,000 RIRs from 260 rooms. Experiments show that our method strongly outperforms a series of baselines. Furthermore, we successfully perform sim-to-real transfer by evaluating our model on four realworld environments, demonstrating the generalizability of our approach and the realism of our dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SonoWorld: From One Image to a 3D Audio-Visual SceneDerong Jin, Xiyi Chen, Ming C. Lin, Ruohan GaoCVPR 2026 · 被引用 4 次
- Resounding Acoustic Fields with ReciprocityZitong Lan, Yiduo Hao, Mingmin ZhaoNeurIPS 2025 · 被引用 3 次
- Few-shot Acoustic Synthesis with Multimodal Flow MatchingAmandine BrunettoCVPR 2026 · 被引用 2 次
- Differentiable Room Acoustic Rendering with Multi-View Vision PriorsDerong Jin, Ruohan GaoICCV 2025
它引用的顶会 Paper23
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 被引用 92 次
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 被引用 80 次
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar 等NeurIPS 2023 · 被引用 77 次
相关 Paper
- Multimodal Neural Acoustic Fields for Immersive Mixed RealityGuaneen Tong, Johnathan Chi-Ho Leung, Xi Peng, Haosheng Shi 等IEEE VR 2025 · 被引用 4 次
- Hearing Anything AnywhereMason Long Wang, Ryosuke Sawata, Samuel Clarke, Ruohan Gao 等CVPR 2024 · 被引用 6 次
- GWA: A Large High-Quality Acoustic Dataset for Audio ProcessingZhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, Dinesh ManochaSIGGRAPH 2022 · 被引用 23 次
- Deep Neural Room Acoustics PrimitiveYuhang He, Anoop Cherian, Gordon Wichern, Andrew MarkhamICML 2024 · 被引用 6 次
- AV-RIR: Audio-Visual Room Impulse Response EstimationAnton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya 等CVPR 2024 · 被引用 15 次
