RAGG: Retrieval-Augmented Grasp Generation Model
Zhenhua Tang, Bin Zhu, Yanbin Hao, Chong-Wah Ngo, Richang Hong
Abstract
Intent-based grasp generation inherently involves challenges such as manipulation ambiguity and modality gaps. To address these, we propose a novel Retrieval-Augmented Grasp Generation model (RAGG). Our key insight is that when humans manipulate new objects, they initially mimic the interaction patterns observed in similar objects, then progressively adjust hand-object contact. Consequently, we develop RAGG as a two-stage approach, encompassing retrieval-guided generation and structurally stable grasp refinement. In the first stage, we propose a Retrieval-Augmented Diffusion Model (ReDim), which identifies the most relevant interaction instance from a knowledge base to explicitly guide grasp generation, thereby mitigating ambiguity and bridging modality gaps to ensure semantically correct manipulation. In the second stage, we introduce a Progressive Refinement Network (PRN) with Kolmogorov-Arnold Network (KAN) layers to refine the generated coarse grasp, employing a Structural Similarity Index loss to constrain the spatial relationship between the hand and the object, thus ensuring the stability of the grasp. Extensive experiments on the OakInk and GRAB benchmarks demonstrate that RAGG achieves superior results compared to the state-of-the-art approach, indicating not only better physical feasibility and controllability, but also strong generalization and interpretability for unseen objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- ReMoDiffuse: Retrieval-Augmented Motion Diffusion ModelMingyuan Zhang, Xinying Guo, Liang Pan, Zhongang Cai et al.ICCV 2023 · 301 citations
- Hand-Object Contact Consistency Reasoning for Human Grasps GenerationHanwen Jiang, Shaowei Liu, Jiashun Wang, Xiaolong WangICCV 2021 · 242 citations
- Retrieval-Augmented Diffusion ModelsAndreas Blattmann, Robin Rombach, Kaan Oktay, Jonas Müller et al.NeurIPS 2022 · 239 citations
Related papers
- Retrieving Semantics from the Deep: an RAG Solution for Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg et al.CVPR 2025
- AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp SynthesisXiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma et al.CVPR 2026 · 1 citation
- RealRAG: Retrieval-augmented Realistic Image Generation via Self-reflective Contrastive LearningYuanhuiyi Lyu, Xu Zheng, Lutao Jiang, Yibo Yan et al.ICML 2025
- OakInk: A Large-scale Knowledge Repository for Understanding Hand-Object InteractionLixin Yang, Kailin Li, Xinyu Zhan, Fei Wu et al.CVPR 2022 · 79 citations
- G-DexGrasp: Generalizable Dexterous Grasping Synthesis via Part-Aware Prior Retrieval and Prior-Assisted GenerationJuntao Jian, Xiuping Liu, Zixuan Chen, Manyi Li et al.ICCV 2025 · 2 citations
