How Would it Sound? Material-Controlled Multimodal Acoustic Profile Generation for Indoor Scenes
Mahnoor Fatima Saad, Ziad Al-Halah
Abstract
How would the sound in a studio change with a carpeted floor and acoustic tiles on the walls? We introduce the task of material-controlled acoustic profile generation, where, given an indoor scene with specific audio-visual characteristics, the goal is to generate a target acoustic profile based on a user-defined material configuration at inference time. We address this task with a novel encoder-decoder approach that encodes the scene's key properties from an audio-visual observation and generates the target Room Impulse Response (RIR) conditioned on the material specifications provided by the user. Our model enables the generation of diverse RIRs based on various material configurations defined dynamically at inference time. To support this task, we create a new benchmark, the Acoustic Wonderland Dataset, designed for developing and evaluating materialaware RIR prediction methods under diverse and challenging settings. Our results demonstrate that the proposed model effectively encodes material information and generates high-fidelity RIRs, outperforming several baselines and state-of-the-art methods. Project: https://mahnoor- fatima-saad.github.io/m-capa.html
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- PAVAS: Physics-Aware Video-to-Audio SynthesisOh Hyun-Bin, Yuhta Takida, Toshimitsu Uesaka, Tae-Hyun Oh et al.CVPR 2026 · 2 citations
- Few-shot Acoustic Synthesis with Multimodal Flow MatchingAmandine BrunettoCVPR 2026 · 2 citations
Builds on16
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans et al.NeurIPS 2021 · 826 citations
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum et al.NeurIPS 2022 · 153 citations
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 92 citations
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 80 citations
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar et al.NeurIPS 2023 · 77 citations
Related papers
- GWA: A Large High-Quality Acoustic Dataset for Audio ProcessingZhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, Dinesh ManochaSIGGRAPH 2022 · 23 citations
- Hearing Anywhere in Any EnvironmentXiulong Liu, Anurag Kumar, Paul Calamia, Sebastià Vicenc Amengual Garí et al.CVPR 2025
- Scene-Aware Audio Rendering via Deep Acoustic AnalysisZhenyu Tang, Nicholas J. Bryan, Dingzeyu Li, Timothy R. Langlois et al.IEEE VR 2020 · 37 citations
- Real Acoustic Fields: An Audio-Visual Room Acoustics Dataset and BenchmarkZiyang Chen, Israel D. Gebru, Christian Richardt, Anurag Kumar et al.CVPR 2024
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 42 citations
