AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence
Jiawei Zhang, Kaizhe Hu, Yingqian Huang, Yuanchen Ju, Zhengrong Xue, Huazhe Xu
Abstract
Despite the recent success of modern imitation learning methods in robot manipulation, their performance is often limited to specific object shapes due to the constrained data diversity. Leveraging powerful 3D generative models and vision foundation models (VFM), the proposed AffordGen framework overcomes this limitation by utilizing the semantic correspondence of meaningful keypoints across large-scale 3D meshes to generate new robot manipulation tra-jectories. This large-scale, affordance-aware dataset is then used to train a robust, closed-loop visuomotor policy, combining the semantic generalizability of affordances with the reactive robustness of end-to-end learning. Experiments in simulation and the real world show that policies trained with AffordGen achieve high success rates and enable zero-shot generalization to truly unseen objects, significantly im-proving data efficiency in robot learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6fd2bbdc-bc89-45a1-8a36-a99f9fb3cd19Builds on7
- A Tale of Two Features: Stable Diffusion Complements DINO for Zero-Shot Semantic CorrespondenceJunyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Polania Cabrera et al.NeurIPS 2023 · 371 citations
- GenSim: Generating Robotic Simulation Tasks via Large Language ModelsLirui Wang, Yiyang Ling, Zhecheng Yuan, Mohit Shridhar et al.ICLR 2024 · 143 citations
- AffordDP: Generalizable Diffusion Policy with Transferable AffordanceShijie Wu, Yihang Zhu, Yunao Huang, Kaizhen Zhu et al.CVPR 2025
- Data Scaling Laws in Imitation Learning for Robotic ManipulationFanqi Lin, Yingdong Hu, Pingyue Sheng, Chuan Wen et al.ICLR 2025
- RDT-1B: a Diffusion Foundation Model for Bimanual ManipulationSongming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan et al.ICLR 2025
Related papers
- VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic ManipulationHanzhi Chen, Boyang Sun, Anran Zhang, Marc Pollefeys et al.CVPR 2025
- UniJEPA: Enhancing Robot Policy via Unified Continuous and Discrete Representation LearningJianke Zhang, Yucheng Hu, Yanjiang Guo, Xiaoyu Chen et al.ICML 2026
- GENMANIP: LLM-driven Simulation for Generalizable Instruction-Following ManipulationNing Gao, Yilun Chen, Shuai Yang, Xinyi Chen et al.CVPR 2025
- Towards Affordance-Aware Robotic Dexterous Grasping with Human-like PriorsHaoyu Zhao, Linghao Zhuang, Xingyue Zhao, Cheng Zeng et al.AAAI 2026 · 4 citations
- Adaptive Articulated Object Manipulation on the Fly with Foundation Model Reasoning and Part GroundingXiaojie Zhang, Yuanfei Wang, Ruihai Wu, Kunqi Xu et al.ICCV 2025 · 2 citations
