Enhancing XR Auditory Realism via Multimodal Scene-Aware Acoustic Rendering
Tianyu Xu, Jihan Li, Penghe Zu, Pranav Sahay, Maruchi Kim, Jack Obeng-Marnu, Farley Miller, Xun Qian, Katrina Passarella, Mahitha Rachumalla, Rajeev Nongpiur, D. Shin
Abstract
In Extended Reality (XR), rendering sound that accurately simulates real-world acoustics is pivotal in creating lifelike and believable virtual experiences. However, existing XR spatial audio rendering methods often struggle with real-time adaptation to diverse physical scenes, causing a sensory mismatch between visual and auditory cues that disrupts user immersion. To address this, we introduce SAMOSA, a novel on-device system that renders spatially accurate sound by dynamically adapting to its physical environment. SAMOSA leverages a synergistic multimodal scene representation by fusing real-time estimations of room geometry, surface materials, and semantic-driven acoustic context. This rich representation then enables efficient acoustic calibration via scene priors, allowing the system to synthesize a highly realistic Room Impulse Response (RIR). We validate our system through technical evaluation using acoustic metrics for RIR synthesis across various room configurations and sound types, alongside an expert evaluation (N=12). Evaluation results demonstrate SAMOSA's feasibility and efficacy in enhancing XR auditory realism.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 974b6cbb-5a4e-41dc-8200-c9c703efa017Cited by top-tier papers2
- Resounding Acoustic Fields with ReciprocityZitong Lan, Yiduo Hao, Mingmin ZhaoNeurIPS 2025 · 3 citations
- MoXaRt: Audio-Visual Object-Guided Sound Interaction for XRTianyu Xu, Sieun Kim, Qianhui Zheng, Ruoyu Xu et al.CHI 2026 · 1 citation
Builds on10
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum et al.NeurIPS 2022 · 153 citations
- DepthLab: Real-time 3D Interaction with Depth Maps for Mobile Augmented RealityRuofei Du, Eric Turner, Maksym Dzitsiuk, Luca Prasso et al.UIST 2020 · 145 citations
- Image2Reverb: Cross-Modal Reverb Impulse Response SynthesisNikhil Singh, Jeff Mentch, Jerry Ng, Matthew Beveridge et al.ICCV 2021 · 61 citations
- Acoustic Volume Rendering for Neural Impulse Response FieldsZitong Lan, Chenhao Zheng, Zhiwei Zheng, Mingmin ZhaoNeurIPS 2024 · 35 citations
Related papers
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 92 citations
- Differentiable Room Acoustic Rendering with Multi-View Vision PriorsDerong Jin, Ruohan GaoICCV 2025
- Hearing Anywhere in Any EnvironmentXiulong Liu, Anurag Kumar, Paul Calamia, Sebastià Vicenc Amengual Garí et al.CVPR 2025
- Auptimize: Optimal Placement of Spatial Audio Cues for Extended RealityHyunsung Cho, Alexander Wang, Divya Kartik, Emily Liying Xie et al.UIST 2024 · 19 citations
- Scene2Hap: Generating Scene-Wide Haptics for VR from Scene Context with Multimodal LLMsArata Jingu, Easa AliAbbasi, Sara Safaee, Paul Strohmeier et al.CHI 2026 · 4 citations
