Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
Ruihai Wu, Kai Cheng, Yan Zhao, Chuanruo Ning, Guanqi Zhan, Hao Dong
摘要
Perceiving and manipulating 3D articulated objects in diverse environments is essential for home-assistant robots. Recent studies have shown that point-level affordance provides actionable priors for downstream manipulation tasks. However, existing works primarily focus on single-object scenarios with homogeneous agents, overlooking the realistic constraints imposed by the environment and the agent's morphology, e.g., occlusions and physical limitations. In this paper, we propose an environment-aware affordance framework that incorporates both object-level actionable priors and environment constraints. Unlike object-centric affordance approaches, learning environment-aware affordance faces the challenge of combinatorial explosion due to the complexity of various occlusions, characterized by their quantities, geometries, positions and poses. To address this and enhance data efficiency, we introduce a novel contrastive affordance learning framework capable of training on scenes containing a single occluder and generalizing to scenes with complex occluder combinations. Experiments demonstrate the effectiveness of our proposed approach in learning affordance considering environment constraints. Introduction Articulated objects, such as doors and drawers, exist everywhere in our daily life. Perceiving and manipulating these objects present crucial yet challenging tasks in computer vision and robotics. Unlike rigid objects, articulated objects exhibit diverse articulation types and functionally important articulated parts crucial for human and robot interactions. Numerous research endeavors have been investigating articulated objects broadly, encompassing joint parameters estimation [41, 47] , part pose estimation [22, 23] , kinematic structure estimation [34, 33] , digital twins generalization [13, 10] , articulated part robotic manipulation [25, 44, 46, 4] and few-shot policy adaptation [42] . However, most existing works for manipulating articulated objects primarily focus on single-object scenarios with homogeneous agents, such as flying grippers [46, 25, 44] or fixed-position robot arms [6] . Consequently, these approaches tend to develop object-centric representations and policies, neglecting the realistic constraints imposed by both the environment and the agent's morphology. These constraints are commonplace in real-world scenarios and their oversight limits the applicability and performance of the manipulation tasks. For example, successfully opening a cabinet door that is obstructed by occluders not only depends on the properties of the target door but also heavily relies on the robot's position and the way it interacts (e.g., colliding or bypassing) with the occluders. We take a significant step towards manipulating articulated objects in a more realistic setting, i.e., considering constraints imposed by the environment and robot. Such a task encounters the combi- * Equal contribution. Author ordering determined by coin flip.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- GarmentLab: A Unified Simulation and Benchmark for Garment ManipulationHaoran Lu, Ruihai Wu, Yitong Li, Sijie Li 等NeurIPS 2024 · 被引用 37 次
- Amodal Ground Truth and Completion in the WildGuanqi Zhan, Chuanxia Zheng, Weidi Xie, Andrew ZissermanCVPR 2024 · 被引用 23 次
- Learning 2D Invariant Affordance Knowledge for 3D Affordance GroundingXianqiang Gao, Pingrui Zhang, Delin Qu, Dong Wang 等AAAI 2025 · 被引用 20 次
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language ModelsXinyi Wang, Xun Yang, Yanlong Xu, Yuchen Wu 等NeurIPS 2025 · 被引用 18 次
- UniGarmentManip: A Unified Framework for Category-Level Garment Manipulation via Dense Visual CorrespondenceRuihai Wu, Haoran Lu, Yiyan Wang, Yubo Wang 等CVPR 2024 · 被引用 14 次
它引用的顶会 Paper15
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta 等ICCV 2021 · 被引用 240 次
- Graspness Discovery in Clutters for Fast and Accurate Grasp DetectionChenxi Wang, Haoshu Fang, Minghao Gou, Hongjie Fang 等ICCV 2021 · 被引用 177 次
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo 等ICLR 2022 · 被引用 119 次
- Ditto: Building Digital Twins of Articulated Objects from InteractionZhenyu Jiang, Cheng-Chun Hsu, Yuke ZhuCVPR 2022 · 被引用 77 次
- Act the Part: Learning Interaction Strategies for Articulated Object Part DiscoverySamir Yitzhak Gadre, Kiana Ehsani, Shuran SongICCV 2021 · 被引用 64 次
相关 Paper
- DualAfford: Learning Collaborative Visual Affordance for Dual-gripper ManipulationYan Zhao, Ruihai Wu, Zhehuan Chen, Yourong Zhang 等ICLR 2023 · 被引用 2 次
- AdaManip: Adaptive Articulated Object Manipulation Environments and Policy LearningYuanfei Wang, Xiaojie Zhang, Ruihai Wu, Yu Li 等ICLR 2025
- Where2Explore: Few-shot Affordance Learning for Unseen Novel Categories of Articulated ObjectsChuanruo Ning, Ruihai Wu, Haoran Lu, Kaichun Mo 等NeurIPS 2023 · 被引用 64 次
- PA3FF: Learning Part-Aware Dense 3D Feature Field For Generalizable Articulated Object ManipulationYue Chen, Muqing Jiang, Kaifeng Zheng, Jiaqi Liang 等ICLR 2026 · 被引用 2 次
- OVA-Fields: Weakly Supervised Open-Vocabulary Affordance Fields for Robot Operational Part DetectionHeng Su, Mengying Xie, Nieqing Cao, Yan Ding 等ICCV 2025 · 被引用 2 次
