ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
Peiming Li, Ziyi Wang, Mengyuan Liu, Hong Liu, Chen Chen
摘要
Grasp generation aims to create complex hand-object interactions with a specified object. While traditional approaches for hand generation have primarily focused on visibility and diversity under scene constraints, they tend to overlook the fine-grained hand-object interactions such as contacts, resulting in inaccurate and undesired grasps. To address these challenges, we propose a controllable grasp generation task and introduce ClickDiff, a controllable conditional generation model that leverages a fine-grained Semantic Contact Map (SCM). Particularly when synthesizing interactive grasps, the method enables the precise control of grasp synthesis through either user-specified or algorithmically predicted Semantic Contact Map. Specifically, to optimally utilize contact supervision constraints and to accurately model the complex physical structure of hands, we propose a Dual Generation Framework. Within this framework, the Semantic Conditional Module generates reasonable contact maps based on fine-grained contact information, while the Contact Conditional Module utilizes contact maps alongside object point clouds to generate realistic grasps. We evaluate the evaluation criteria applicable to controllable grasp generation. Both unimanual and bimanual generation experiments on GRAB and ARCTIC datasets verify the validity of our proposed method, demonstrating the efficacy and robustness of ClickDiff, even with previously unseen objects. Our code is available at https://github.com/adventurer-w/ClickDiff.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language ModelZhenhao Zhang, Ye Shi, Lingxiao Yang, Suting Ni 等NeurIPS 2025 · 被引用 25 次
- Inversion-DPO: Precise and Efficient Post-Training for Diffusion ModelsZejian Li, Yize Li, Chenye Meng, Zhongni Liu 等ACM MM 2025
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Planning with Diffusion for Flexible Behavior SynthesisMichael Janner, Yilun Du, Joshua B. Tenenbaum, Sergey LevineICML 2022 · 被引用 1,115 次
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++Ruilong Li, Shan Yang, David A. Ross, Angjoo KanazawaICCV 2021 · 被引用 701 次
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 被引用 486 次
- PhysDiff: Physics-Guided Human Motion Diffusion ModelYe Yuan, Jiaming Song, Umar Iqbal, Arash Vahdat 等ICCV 2023 · 被引用 414 次
相关 Paper
- ContactGen: Generative Contact Modeling for Grasp GenerationShaowei Liu, Yang Zhou, Jimei Yang, Saurabh Gupta 等ICCV 2023 · 被引用 60 次
- Task-Oriented Human Grasp Synthesis via Context- and Task-Aware DiffusersAn-Lun Liu, Yu-Wei Chao, Yi-Ting ChenICCV 2025 · 被引用 1 次
- TriDi: Trilateral Diffusion of 3D Humans, Objects, and InteractionsIlia A. Petrov, Riccardo Marin, Julian Chibane, Gerard Pons-MollICCV 2025 · 被引用 1 次
- Contact Map Transfer with Conditional Diffusion Model for Generalizable Dexterous Grasp GenerationYiyao Ma, Kai Chen, Kexin Zheng, Qi DouNeurIPS 2025 · 被引用 6 次
- AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp SynthesisXiaofei Wu, Yi Zhang, Yumeng Liu, Yuexin Ma 等CVPR 2026 · 被引用 1 次
