HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models
Khalequzzaman Chowdhury Sayem, Mubarrat Chowdhury, Yihalem Yimolal Tiruneh, Muneeb Ahmed Khan, Muhammad Salman Ali, Binod Bhattarai, Seungryul Baek
摘要
Understanding the fine-grained articulation of human hands is critical in high-stakes settings such as robot-assisted surgery, chip manufacturing, and AR/VR-based human–AI interaction. Despite achieving near-human performance on general vision-language benchmarks, current vision-language models (VLMs) struggle with fine-grained spatial reasoning—especially in interpreting complex, articulated hand poses. We introduce HandVQA, a large-scale diagnostic benchmark designed to evaluate VLMs' understanding of detailed hand anatomy through visual question answering. Built upon high-quality 3D hand datasets (FreiHAND, InterHand2.6M, FPHA), our benchmark includes over 1.6M controlled multiple-choice questions that probe spatial relationships between hand joints, such as angles, distances, and relative positions. We evaluate several state-of-the-art VLMs (LLaVA, DeepSeek and Qwen-VL) in both base and fine-tuned settings, using lightweight fine-tuning via LoRA. Our findings reveal systematic limitations in current models, including hallucinated finger parts, incorrect geometric interpretations, and poor generalization. HandVQA not only exposes these critical reasoning gaps but provides a validated path to improvement. We demonstrate that the 3D-grounded spatial knowledge learned from our benchmark transfers in a zero-shot setting, significantly improving accuracy of model on novel downstream tasks like hand gesture recognition (+10.33%) and hand-object interaction (+2.63%). Code and dataset will be released upon acceptance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- FreiHAND: A Dataset for Markerless Capture of Hand Pose and Shape From Single RGB ImagesChristian Zimmermann, Duygu Ceylan, Jimei Yang, Bryan C. Russell 等ICCV 2019 · 被引用 493 次
相关 Paper
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction DynamicsMasatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato 等CVPR 2026 · 被引用 2 次
- BOP-ASK: Object-Interaction Reasoning for Vision-Language ModelsVineet Bhat, Sungsu Kim, Valts Blukis, Greg Heinrich 等CVPR 2026 · 被引用 6 次
- Do 3D Large Language Models Really Understand 3D Spatial Relationships?Xianzheng Ma, Tao Sun, Shuai Chen, Yash Bhalgat 等ICLR 2026 · 被引用 7 次
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han 等NeurIPS 2025 · 被引用 159 次
- FloorplanQA: A Benchmark for Spatial Reasoning in LLMs using Structured RepresentationsFedor Rodionov, Abdelrahman Eldesokey, Michael Birsak, John Femiani 等ICML 2026 · 被引用 14 次
