Audio-Visual Floorplan Reconstruction
Senthil Purushwalkam, Sebastià Vicenc Amengual Garí, Vamsi Krishna Ithapu, Carl Schissler, Philip W. Robinson, Abhinav Gupta, Kristen Grauman
摘要
Given only a few glimpses of an environment, how much can we infer about its entire floorplan? Existing methods can map only what is visible or immediately apparent from context, and thus require substantial movements through a space to fully map it. We explore how both audio and visual sensing together can provide rapid floorplan reconstruction from limited viewpoints. Audio not only helps sense geometry outside the camera’s field of view, but it also reveals the existence of distant freespace (e.g., a dog barking in another room) and suggests the presence of rooms not visible to the camera (e.g., a dishwasher humming in what must be the kitchen to the left). We introduce AV-Map, a novel multi-modal encoder-decoder framework that reasons jointly about audio and vision to reconstruct a floorplan from a short input video sequence. We train our model to predict both the interior structure of the environment and the associated rooms’ semantic labels. Our results on 85 large real-world environments show the impact: with just a few glimpses spanning 26% of an area, we can estimate the whole area with 66% accuracy—substantially better than the state of the art approach for extrapolating visual maps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Pathdreamer: A World Model for Indoor NavigationJing Yu Koh, Honglak Lee, Yinfei Yang, Jason Baldridge 等ICCV 2021 · 被引用 128 次
- Few-Shot Audio-Visual Learning of Environment AcousticsSagnik Majumder, Changan Chen, Ziad Al-Halah, Kristen GraumanNeurIPS 2022 · 被引用 80 次
- Move2Hear: Active Audio-Visual Source SeparationSagnik Majumder, Ziad Al-Halah, Kristen GraumanICCV 2021 · 被引用 48 次
- Disentangled Counterfactual Learning for Physical Audiovisual Commonsense ReasoningChangsheng Lv, Shuai Zhang, Yapeng Tian, Mengshi Qi 等NeurIPS 2023 · 被引用 26 次
- GWA: A Large High-Quality Acoustic Dataset for Audio ProcessingZhenyu Tang, Rohith Aralikatti, Anton Jeran Ratnarajah, Dinesh ManochaSIGGRAPH 2022 · 被引用 23 次
它引用的顶会 Paper5
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta 等ICLR 2020 · 被引用 603 次
- Self-Supervised Moving Vehicle Tracking With Stereo SoundChuang Gan, Hang Zhao, Peihao Chen, David D. Cox 等ICCV 2019 · 被引用 157 次
- Floor-SP: Inverse CAD for Floorplans by Sequential Room-Wise Shortest PathJiacheng Chen, Chen Liu, Jiaye Wu, Yasutaka FurukawaICCV 2019 · 被引用 90 次
- Learning to Set Waypoints for Audio-Visual NavigationChangan Chen, Sagnik Majumder, Ziad Al-Halah, Ruohan Gao 等ICLR 2021 · 被引用 28 次
相关 Paper
- Chat2Map: Efficient Scene Mapping from Multi-Ego ConversationsSagnik Majumder, Hao Jiang, Pierre Moulon, Ethan Henderson 等CVPR 2023
- AVLEN: Audio-Visual-Language Embodied Navigation in 3D EnvironmentsSudipta Paul, Amit Roy-Chowdhury, Anoop CherianNeurIPS 2022 · 被引用 43 次
- Semantic Audio-Visual NavigationChangan Chen, Ziad Al-Halah, Kristen GraumanCVPR 2021
- Dense 2D-3D Indoor Prediction with Sound via Aligned Cross-Modal DistillationHeeseung Yun, Joonil Na, Gunhee KimICCV 2023 · 被引用 8 次
- Hear you are: Teaching LLMs Spatial Reasoning with Vision and Spatial SoundHyeonggon Ryu, Joon Son Chung, David HarwathCVPR 2026 · 被引用 4 次
