Be Everywhere - Hear Everything (BEE): Audio Scene Reconstruction by Sparse Audio-Visual Samples
Mingfei Chen, Kun Su, Eli Shlizerman
摘要
Fully immersive and interactive audio-visual scenes are dynamic such that the listeners and the sound emitters move and interact with each other. Reconstruction of an immersive sound experience, as it happens in the scene, requires detailed reconstruction of the audio perceived by the listener at an arbitrary location. The audio at the listener location is a complex outcome of sound propagation through the scene geometry and interacting with surfaces and also the locations of the emitters and the sounds they emit. Due to these aspects, detailed audio reconstruction requires extensive sampling of audio at any potential listener location. This is usually difficult to implement in realistic real-time dynamic scenes. In this work, we propose to circumvent the need for extensive sensors by leveraging audio and visual samples from only a handful of A/V receivers placed in the scene. In particular, we introduce a novel method and end-to-end integrated rendering pipeline which allows the listener to be everywhere and hear everything (BEE) in a dynamic scene in real-time. BEE reconstructs the audio with two main modules, Joint Audio-Visual Representation, and Integrated Rendering Head. The first module extracts the informative audio-visual features of the scene from sparse A/V reference samples, while the second module integrates the audio samples with learned time-frequency transformations to obtain the target sound. Our experiments indicate that BEE outperforms existing methods by a large margin in terms of quality of sound reconstruction, can generalize to scenes not seen in training and runs in real-time speed.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 被引用 22 次
- AV-RIR: Audio-Visual Room Impulse Response EstimationAnton Ratnarajah, Sreyan Ghosh, Sonal Kumar, Purva Chiniya 等CVPR 2024 · 被引用 15 次
- Resounding Acoustic Fields with ReciprocityZitong Lan, Yiduo Hao, Mingmin ZhaoNeurIPS 2025 · 被引用 3 次
- SoundVista: Novel-View Ambient Sound Synthesis via Visual-Acoustic BindingMingfei Chen, Israel D. Gebru, Ishwarya Ananthabhotla, Christian Richardt 等CVPR 2025
- Video-Guided Foley Sound Generation with Multimodal ControlsZiyang Chen, Prem Seetharaman, Bryan C. Russell, Oriol Nieto 等CVPR 2025
它引用的顶会 Paper5
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Learning Neural Acoustic FieldsAndrew F. Luo, Yilun Du, Michael J. Tarr, Josh Tenenbaum 等NeurIPS 2022 · 被引用 153 次
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 被引用 80 次
- Neural Synthesis of Binaural Speech From Mono AudioAlexander Richard, Dejan Markovic, Israel D. Gebru, Steven Krenn 等ICLR 2021 · 被引用 73 次
- Visual Acoustic MatchingChangan Chen, Ruohan Gao, Paul Calamia, Kristen GraumanCVPR 2022 · 被引用 42 次
相关 Paper
- AV-Cloud: Spatial Audio Rendering Through Audio-Visual Cloud SplattingMingfei Chen, Eli ShlizermanNeurIPS 2024 · 被引用 14 次
- Hearing Anything AnywhereMason Long Wang, Ryosuke Sawata, Samuel Clarke, Ruohan Gao 等CVPR 2024 · 被引用 6 次
- INRAS: Implicit Neural Representation for Audio ScenesKun Su, Mingfei Chen, Eli ShlizermanNeurIPS 2022 · 被引用 92 次
- Directional sources and listeners in interactive sound propagation using reciprocal wave field codingChakravarty R. Alla Chaitanya, Nikunj Raghuvanshi, Keith W. Godin, Zechen Zhang 等SIGGRAPH 2020 · 被引用 16 次
- AV-NeRF: Learning Neural Fields for Real-World Audio-Visual Scene SynthesisSusan Liang, Chao Huang, Yapeng Tian, Anurag Kumar 等NeurIPS 2023 · 被引用 77 次
