Motion Question Answering via Modular Motion Programs
Mark Endo, Joy Hsu, Jiaman Li, Jiajun Wu
摘要
In order to build artificial intelligence systems that can perceive and reason with human behavior in the real world, we must first design models that conduct complex spatio-temporal reasoning over motion sequences. Moving towards this goal, we propose the HumanMotionQA task to evaluate complex, multi-step reasoning abilities of models on long-form human motion sequences. We generate a dataset of question-answer pairs that require detecting motor cues in small portions of motion sequences, reasoning temporally about when events occur, and querying specific motion attributes. In addition, we propose NSPose, a neuro-symbolic method for this task that uses symbolic reasoning and a modular design to ground motion through learning motion concepts, attribute neural operators, and temporal relations. We demonstrate the suitability of NSPose for the HumanMotionQA task, outperforming all baseline methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- What's Left? Concept Grounding with Logic-Enhanced Foundation ModelsJoy Hsu, Jiayuan Mao, Joshua B. Tenenbaum, Jiajun WuNeurIPS 2023 · 被引用 54 次
- Iterative Motion Editing with Natural LanguagePurvi Goel, Kuan-Chieh Wang, C. Karen Liu, Kayvon FatahalianSIGGRAPH 2024 · 被引用 22 次
- LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive TokensZekun Li, Sizhe An, Chengcheng Tang, Chuan Guo 等CVPR 2026 · 被引用 12 次
- Weakly-Supervised Emotion Transition Learning for Diverse 3D Co-Speech Gesture GenerationXingqun Qi, Jiahao Pan, Peng Li, Ruibin Yuan 等CVPR 2024 · 被引用 9 次
- How Can Objects Help Video-Language Understanding?Zitian Tang, Shijie Wang, Junho Cho, Jaewook Yoo 等ICCV 2025 · 被引用 8 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll 等ICCV 2019 · 被引用 1,784 次
- Channel-wise Topology Refinement Graph Convolution for Skeleton-Based Action RecognitionYuxin Chen, Ziqi Zhang, Chunfeng Yuan, Bing Li 等ICCV 2021 · 被引用 871 次
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin 等CVPR 2022 · 被引用 752 次
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
相关 Paper
- Image Manipulation via Multi-Hop Instructions - A New Dataset and Weakly-Supervised Neuro-Symbolic ApproachHarman Singh, Poorva Garg, Mohit Gupta, Kevin Shah 等EMNLP 2023
- NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic ReasoningSahil Shah, S. P. Sharan, Harsh Goel, Minkyu Choi 等AAAI 2026 · 被引用 4 次
- Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel LevelAndong Deng, Tongjia Chen, Shoubin Yu, Taojiannan Yang 等CVPR 2025
- Hierarchical Motion Understanding via Motion ProgramsSumith Kulal, Jiayuan Mao, Alex Aiken, Jiajun WuCVPR 2021
- An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue GenerationShiquan Yang, Rui Zhang, Sarah M. Erfani, Jey Han LauACL 2022 · 被引用 17 次
