Imore: Implicit Program-Guided Reasoning for Human Motion QA
Chen Li, Chinthani Sugandhika, Ee Yeo Keat, Eric P. Xing, Hao Zhang, Hong Yang, Deepu Rajan, Basura Fernando
摘要
Existing human motion methods rely on explicit program execution, where the requirement for manually defined functional modules may limit the scalability and adaptability. To overcome this, we propose an implicit program-guided motion reasoning (IMoRe) framework that unifies reasoning across multiple query types without manually designed modules. Unlike existing implicit reasoning approaches that infer reasoning operations from question words, our model directly conditions on structured program functions, ensuring a more precise execution of reasoning steps. Additionally, we introduce a program-guided reading mechanism, which dynamically selects multi-level motion representations from a pretrained motion Vision Transformer (ViT), capturing both high-level semantics and fine-grained motion cues. The reasoning module iteratively refines memory representations, leveraging structured program functions to extract relevant information for different query types. Our model achieves state-of-the-art performance on Babel-QA and generalizes to a newly constructed motion Q&A dataset based on HuMMan, demonstrating its adaptability across different motion reasoning datasets. Code and dataset are available at https://github.com/LUNAProject22/IMoRe.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive TokensZekun Li, Sizhe An, Chengcheng Tang, Chuan Guo 等CVPR 2026 · 被引用 12 次
- PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module LearningThanh-Son Nguyen, Hong Yang, Tzeh Yuan Neoh, Hao Zhang 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- FineMoGen: Fine-Grained Spatio-Temporal Motion Generation and EditingMingyuan Zhang, Huirong Li, Zhongang Cai, Jiawei Ren 等NeurIPS 2023 · 被引用 132 次
相关 Paper
- Question Aware Vision Transformer for Multimodal ReasoningRoy Ganz, Yair Kittenplon, Aviad Aberdam, Elad Ben-Avraham 等CVPR 2024
- A Unified End-to-End Retriever-Reader Framework for Knowledge-based VQAYangyang Guo, Liqiang Nie, Yongkang Wong, Yibing Liu 等ACM MM 2022 · 被引用 41 次
- Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelAndong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang 等CVPR 2025
- Exploring Vision Transformers for 3D Human Motion-Language Models with Motion PatchesQing Yu, Mikihiro Tanaka, Kent FujiwaraCVPR 2024 · 被引用 5 次
- ViHOI: Human-Object Interaction Synthesis with Visual PriorsSongjin Cai, Linjie Zhong, Ling Guo, Changxing DingCVPR 2026 · 被引用 2 次
