DualAfford: Learning Collaborative Visual Affordance for Dual-gripper Manipulation
Yan Zhao, Ruihai Wu, Zhehuan Chen, Yourong Zhang, Qingnan Fan, Kaichun Mo, Hao Dong
Abstract
It is essential yet challenging for future home-assistant robots to understand and manipulate diverse 3D objects in daily human environments. Towards building scalable systems that can perform diverse manipulation tasks over various 3D shapes, recent works have advocated and demonstrated promising results learning visual actionable affordance, which labels every point over the input 3D geometry with an action likelihood of accomplishing the downstream task (e.g., pushing or picking-up). However, these works only studied single-gripper manipulation tasks, yet many real-world tasks require two hands to achieve collaboratively. In this work, we propose a novel learning framework, DualAfford, to learn collaborative affordance for dual-gripper manipulation tasks. The core design of the approach is to reduce the quadratic problem for two grippers into two disentangled yet interconnected subtasks for efficient learning. Using the large-scale PartNet-Mobility and ShapeNet datasets, we set up four benchmark tasks for dual-gripper manipulation. Experiments prove the effectiveness and superiority of our method over three baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 44a15547-ff0d-4df3-bdd1-66dc90617705Cited by top-tier papers8
- LEMON: Learning 3D Human-Object Interaction Relation from 2D ImagesYuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao et al.CVPR 2024 · 12 citations
- Seeing the Unseen: Visual Common Sense for Semantic PlacementRam Ramrakhya, Aniruddha Kembhavi, Dhruv Batra, Zsolt Kira et al.CVPR 2024 · 3 citations
- VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic ManipulationHuayi Zhou, Kui JiaICLR 2026 · 3 citations
- BiPreManip: Learning Affordance-Based Bimanual Preparatory Manipulation through Anticipatory CollaborationYan Shen, Feng Jiang, Zichen He, Xiaoqi Li et al.CVPR 2026 · 3 citations
- PA3FF: Learning Part-Aware Dense 3D Feature Field For Generalizable Articulated Object ManipulationYue Chen, Muqing Jiang, Kaifeng Zheng, Jiaqi Liang et al.ICLR 2026 · 2 citations
Builds on9
- 6-DOF GraspNet: Variational Grasp Generation for Object ManipulationArsalan Mousavian, Clemens Eppner, Dieter FoxICCV 2019 · 673 citations
- Hand-Object Contact Consistency Reasoning for Human Grasps GenerationHanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong WangICCV 2021 · 242 citations
- Where2Act: From Pixels to Actions for Articulated 3D ObjectsKaichun Mo, Leonidas J. Guibas, Mustafa Mukadam, Abhinav Gupta et al.ICCV 2021 · 240 citations
- VAT-Mart: Learning Visual Action Trajectory Proposals for Manipulating 3D ARTiculated ObjectsRuihai Wu, Yan Zhao, Kaichun Mo, Zizheng Guo et al.ICLR 2022 · 119 citations
- Deep Imitation Learning for Bimanual Robotic ManipulationFan Xie, Alexander Chowdhury, M. Clara De Paolis Kaluza, Linfeng Zhao et al.NeurIPS 2020 · 108 citations
Related papers
- 3D AffordanceNet: A Benchmark for Visual Object Affordance UnderstandingShengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen et al.CVPR 2021
- MAAL: Multimodality-Aware Autoencoder-based Affordance Learning for 3D Articulated ObjectsYuanzhi Liang, Xiaohan Wang, Linchao Zhu, Yi YangICCV 2023 · 3 citations
- A3D: Adaptive Affordance Assembly with Dual-Arm ManipulationJiaqi Liang, Yue Chen, Qize Yu, Yan Shen et al.AAAI 2026 · 3 citations
- Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under OcclusionsRuihai Wu, Kai Cheng, Yan Zhao, Chuanruo Ning et al.NeurIPS 2023 · 43 citations
- InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable ObjectsXinhao Cai, Minghang Zheng, Xin Jin, Yang LiuACM MM 2025 · 1 citation
