HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato, Takuma Yagi
Abstract
Hand–object interaction (HOI) inherently involves dynamics where human manipulations produce distinct spatio-temporal effects on objects.However, existing semantic HOI benchmarks focused either on manipulation or on the resulting effects at a coarse level, lacking fine-grained spatio-temporal reasoning to capture the underlying dynamics in HOI.We introduce HanDyVQA, a fine-grained video question-answering benchmark that comprehensively covers both the manipulation and effect aspects of HOI. HanDyVQA comprises six complementary question types (Action, Process, Objects, Location, State Change, and Object Parts), totalling 11.1K multiple-choice QA pairs.Collected QA pairs recognizing manipulation styles, hand/object motions, and part-level state changes.HanDyVQA also includes 10.3K segmentation masks for Objects and Object Parts questions, enabling the evaluation of object/part-level reasoning in video object segmentation.We evaluated recent video foundation models on our benchmark and found that even the best-performing model, Gemini-2.5-Pro, reached only 73% average accuracy, which is far from human performance (97%). Further analysis shows the remaining challenges in spatial relationship, motion, and part-level geometric understanding.We also found that integrating explicit HOI-related cues into visual features improves performance, offering insights for developing future models with a deeper understanding of HOI dynamics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3f7c9273-1b51-49aa-bb85-9fcd625b3b4dBuilds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Ego4D: Around the World in 3, 000 Hours of Egocentric VideoKristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis et al.CVPR 2022 · 525 citations
- Detecting Moments and Highlights in Videos via Natural Language QueriesJie Lei, Tamara L. Berg, Mohit BansalNeurIPS 2021 · 425 citations
Related papers
- HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language ModelsKhalequzzaman Chowdhury Sayem, Mubarrat Chowdhury, Yihalem Yimolal Tiruneh, Muneeb Ahmed Khan et al.CVPR 2026 · 3 citations
- Interacted Object Grounding in Spatio-Temporal Human-Object InteractionsXiaoyang Liu, Boran Wen, Xinpeng Liu, Zizheng Zhou et al.AAAI 2025 · 6 citations
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal ModelsAndong Deng, Taojiannan Yang, Shoubin Yu, Lincoln Spencer et al.ICML 2026 · 7 citations
- AGQA: A Benchmark for Compositional Spatio-Temporal ReasoningMadeleine Grunde-McLaughlin, Ranjay Krishna, Maneesh AgrawalaCVPR 2021
- MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingYilun Zhao, Haowei Zhang, Lujing Xie, Tongyan Hu et al.CVPR 2025
