Embodied Amodal Recognition: Learning to Move to Perceive Objects
Jianwei Yang, Zhile Ren, Mingze Xu, Xinlei Chen, David J. Crandall, Devi Parikh, Dhruv Batra
Abstract
Passive visual systems typically fail to recognize objects in the amodal setting where they are heavily occluded. In contrast, humans and other embodied agents have the ability to move in the environment and actively control the viewing angle to better understand object shapes and semantics. In this work, we introduce the task of Embodied Amodel Recognition (EAR): an agent is instantiated in a 3D environment close to an occluded target object, and is free to move in the environment to perform object classification, amodal object localization, and amodal object segmentation. To address this problem, we develop a new model called Embodied Mask R-CNN for agents to learn to move strategically to improve their visual recognition abilities. We conduct experiments using a simulator for indoor environments. Experimental results show that: 1) agents with embodiment (movement) achieve better visual recognition performance than passive ones and 2) in order to improve visual recognition abilities, agents can learn strategic paths that are different from shortest paths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a7d0a0e-775d-4133-a5fa-b09d757ec2b8Cited by top-tier papers22
- Bird's-Eye-View Scene Graph for Vision-Language NavigationRui Liu, Xiaohan Wang, Wenguan Wang, Yi YangICCV 2023 · 100 citations
- SEAL: Self-supervised Embodied Active Learning using Exploration and 3D ConsistencyDevendra Singh Chaplot, Murtaza Dalal, Saurabh Gupta, Jitendra Malik et al.NeurIPS 2021 · 100 citations
- Semantic MapNet: Building Allocentric Semantic Maps and Representations from Egocentric ViewsVincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee et al.AAAI 2021 · 89 citations
- Amodal Segmentation Based on Visible Region Segmentation and Shape PriorYuting Xiao, Yanyu Xu, Ziming Zhong, Weixin Luo et al.AAAI 2021 · 76 citations
- Act the Part: Learning Interaction Strategies for Articulated Object Part DiscoverySamir Yitzhak Gadre, Kiana Ehsani, Shuran SongICCV 2021 · 64 citations
Related papers
- Embodied Visual Active Learning for Semantic SegmentationDavid Nilsson, Aleksis Pirinen, Erik Gärtner, Cristian SminchisescuAAAI 2021 · 37 citations
- Explore and Tell: Embodied Visual Captioning in 3D EnvironmentsAnwen Hu, Shizhe Chen, Liang Zhang, Qin JinICCV 2023 · 4 citations
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical EnvironmentsWeijie Zhou, Xuantang Xiong, Yi Peng, Manli Tao et al.NeurIPS 2025 · 4 citations
- Navigating to Objects Specified by ImagesJacob Krantz, Théophile Gervet, Karmesh Yadav, Austin S. Wang et al.ICCV 2023 · 70 citations
- Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied NavigationZiyu Zhu, Xilin Wang, Yixuan Li, Zhuofan Zhang et al.ICCV 2025 · 11 citations
