Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric Views
Vincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee, Irfan Essa, Dhruv Batra
摘要
We study the task of semantic mapping – specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (‘what is where?’) from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space – navigating to objects seen during the tour (‘Find chair’) or answering questions about the space (‘How many chairs did you see in the house?’). Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length×width×feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01−16.81% (absolute) on mean-IoU and 3.81−19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering. Project page: https://vincentcartillier.github.io/smnet.html.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- MultiON: Benchmarking Semantic Map Memory using Multi-Object NavigationSaim Wani, Shivansh Patel, Unnat Jain, Angel X. Chang 等NeurIPS 2020 · 被引用 156 次
- Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language NavigationPeihao Chen, Dongyu Ji, Kunyang Lin, Runhao Zeng 等NeurIPS 2022 · 被引用 143 次
- Auxiliary Tasks and Exploration Enable ObjectGoal NavigationJoel Ye, Dhruv Batra, Abhishek Das, Erik WijmansICCV 2021 · 被引用 137 次
- GridMM: Grid Memory Map for Vision-and-Language NavigationZihan Wang, Xiangyang Li, Jiahao Yang, Yeqi Liu 等ICCV 2023 · 被引用 136 次
- PONI: Potential Functions for ObjectGoal Navigation with Interaction-free LearningSanthosh Kumar Ramakrishnan, Devendra Singh Chaplot, Ziad Al-Halah, Jitendra Malik 等CVPR 2022 · 被引用 129 次
它引用的顶会 Paper5
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Object Goal Navigation using Goal-Oriented Semantic ExplorationDevendra Singh Chaplot, Dhiraj Gandhi, Abhinav Gupta, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 857 次
- Learning To Explore Using Active Neural SLAMDevendra Singh Chaplot, Dhiraj Gandhi, Saurabh Gupta, Abhinav Gupta 等ICLR 2020 · 被引用 603 次
- Embodied Amodal Recognition: Learning to Move to Perceive ObjectsJianwei Yang, Zhile Ren, Mingze Xu, Xinlei Chen 等ICCV 2019 · 被引用 70 次
- Ego-Topo: Environment Affordances From Egocentric VideoTushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, Kristen GraumanCVPR 2020
相关 Paper
- Episodic Memory Question AnsweringSamyak Datta, Sameer Dharur, Vincent Cartillier, Ruta Desai 等CVPR 2022 · 被引用 23 次
- Learning Navigational Visual Representations with Semantic Map SupervisionYicong Hong, Yang Zhou, Ruiyi Zhang, Franck Dernoncourt 等ICCV 2023 · 被引用 56 次
- MapNav: A Novel Memory Representation via Annotated Semantic Maps for VLM-based Vision-and-Language NavigationLingfeng Zhang, Xiaoshuai Hao, Qinwen Xu, Qiang Zhang 等ACL 2025 · 被引用 55 次
- End-to-End Egospheric Spatial MemoryDaniel James Lenton, Stephen James, Ronald Clark, Andrew J. DavisonICLR 2021 · 被引用 5 次
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 被引用 37 次
