MonSTeR: A Unified Model for Motion, Scene, Text Retrieval
Luca Collorone, Matteo Gioia, Massimiliano Pappa, Paolo Leoni, Giovanni Ficarra, Or Litany, Indro Spinelli, Fabio Galasso
摘要
Intention drives human movement in complex environments, but such movement can only happen if the surrounding context supports it. Despite the intuitive nature of this mechanism, existing research has not yet provided tools to evaluate the alignment between skeletal movement (motion), intention (text), and the surrounding context (scene).
In this work, we introduce MonSTeR, the first MOtioN-Scene-TExt Retrieval model. Inspired by the modeling of higher-order relations, MonSTeR constructs a unified latent space by leveraging unimodal and cross-modal representations. This allows MonSTeR to capture the intricate dependencies between modalities, enabling flexible but robust retrieval across various tasks.
Our results show that MonSTeR outperforms trimodal models that rely solely on unimodal representations. Furthermore, we validate the alignment of our retrieval scores with human preferences through a dedicated user study. We demonstrate the versatility of MonSTeR's latent space on zero-shot in-Scene Object Placement and Motion Captioning. Code and pre-trained models are available at github.com/colloroneluca/MonSTeR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang 等CVPR 2022 · 被引用 462 次
- Weisfeiler and Lehman Go Topological: Message Passing Simplicial NetworksCristian Bodnar, Fabrizio Frasca, Yuguang Wang, Nina Otter 等ICML 2021 · 被引用 315 次
- 3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentZiyu Zhu, Xiaojian Ma, Yixin Chen, Zhidong Deng 等ICCV 2023 · 被引用 247 次
- HUMANISE: Language-conditioned Human Motion Generation in 3D ScenesZan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu 等NeurIPS 2022 · 被引用 207 次
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 被引用 192 次
相关 Paper
- Towards Unified Human Motion-Language Understanding via Sparse Interpretable CharacterizationGuangtao Lyu, Chenghao Xu, Jiexi Yan, Muli Yang 等ICLR 2025
- MotionBind: Multi-Modal Human Motion Alignment for Retrieval, Recognition, and GenerationKaleab Alemayehu Kinfu, René VidalNeurIPS 2025 · 被引用 4 次
- X-MoGen: Unified Motion Generation Across Humans and AnimalsXuan Wang, Kai Ruan, Liyang Qian, Guo Zhi Zhi 等AAAI 2026
- SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action RecognitionNing Wang, Tieyue Wu, Naeha Sharif, Farid Boussaïd 等CVPR 2026 · 被引用 3 次
- ReMoGPT: Part-Level Retrieval-Augmented Motion-Language ModelsQing Yu, Mikihiro Tanaka, Kent FujiwaraAAAI 2025 · 被引用 6 次
