CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wild
Balamurugan Thambiraja, Omid Taheri, Radek Danecek, Giorgio Becherini, Gerard Pons-Moll, Justus Thies
摘要
Hands play a central role in daily life, yet modeling natural hand motions remains underexplored. Existing methods that tackle text-to-hand-motion generation or hand animation captioning rely on studio-captured datasets with limited actions and contexts, making them costly to scale to “in-the-wild” settings. Further, contemporary models and their training schemes struggle to capture animation fidelity with text–motion alignment. To address this, we (1) introduce ‘3D Hands in the Wild’ (3D-HIW), a dataset of 32K 3D hand-motion sequences and aligned text, and (2) propose CLUTCH, an LLM-based hand animation system with two critical innovations: (a) SHIFT, a novel VQ-VAE architecture to tokenize hand motion, and (b) a geometric refinement stage to finetune the LLM. To build 3D- HIW, we propose a data annotation pipeline that combines vision–language models (VLMs) and state-of-the-art 3D hand trackers, and apply it to a large corpus of egocentric action videos covering a wide range of scenarios. To fully capture motion in-the-wild, CLUTCH employs SHIFT, a part–modality decomposed VQ- VAE, which improves generalization and reconstruction fidelity. Finally, to improve animation quality, we introduce a geometric refinement stage, where CLUTCH is co-supervised with a reconstruction loss applied directly to decoded hand motion parameters. Experiments demonstrate state-of-the-art performance on text-to- motion and motion-to-text tasks, establishing the first benchmark for scalable in-the-wild hand motion modelling. Code, data and models will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper27
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu 等NeurIPS 2023 · 被引用 698 次
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 被引用 371 次
- H2O: Two Hands Manipulating Objects for First Person Interaction RecognitionTaein Kwon, Bugra Tekin, Jan Stühmer, Federica Bogo 等ICCV 2021 · 被引用 271 次
- Guided Motion Diffusion for Controllable Human Motion SynthesisKorrawe Karunratanakul, Konpat Preechakul, Supasorn Suwajanakorn, Siyu TangICCV 2023 · 被引用 240 次
相关 Paper
- HOIGPT: Learning Long-Sequence Hand-Object Interaction with Language ModelsMingzhen Huang, Fu-Jen Chu, Bugra Tekin, Kevin J. Liang 等CVPR 2025
- Text-Driven 3D Hand Motion Generation from Sign Language DataLéore Bensabath, Mathis Petrovich, Gül VarolCVPR 2026 · 被引用 5 次
- SHOW3D: Capturing Scenes of 3D Hands and Objects in the WildPatrick Rim, Kevin Harris, Braden Copple, Shangchen Han 等CVPR 2026 · 被引用 5 次
- EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context LearningBinzhu Xie, Shi Qiu, Sicheng Zhang, Yinqiao Wang 等ICLR 2026 · 被引用 4 次
- InteractVLM: 3D Interaction Reasoning from 2D Foundational ModelsSai Kumar Dwivedi, Dimitrije Antic, Shashank Tripathi, Omid Taheri 等CVPR 2025
