TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion Synthesis
Mathis Petrovich, Michael J. Black, Gül Varol
摘要
In this paper, we present TMR, a simple yet effective approach for text to 3D human motion retrieval. While previous work has only treated retrieval as a proxy evaluation metric, we tackle it as a standalone task. Our method extends the state-of-the-art text-to-motion synthesis model TEMOS, and incorporates a contrastive loss to better structure the cross-modal latent space. We show that maintaining the motion generation loss, along with the contrastive training, is crucial to obtain good performance. We introduce a benchmark for evaluation and provide an in-depth analysis by reporting results on several protocols. Our extensive experiments on the KIT-ML and HumanML3D datasets show that TMR outperforms the prior work by a significant margin, for example reducing the median rank from 54 to 19. Finally, we showcase the potential of our approach on moment retrieval. Our code and models are publicly available at https://mathis.petrovich.fr/tmr.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper86
- OmniControl: Control Any Joint at Any Time for Human Motion GenerationYiming Xie, Varun Jampani, Lei Zhong, Deqing Sun 等ICLR 2024 · 被引用 228 次
- HumanTOMATO: Text-aligned Whole-body Motion GenerationShunlin Lu, Ling-Hao Chen, Ailing Zeng, Jing Lin 等ICML 2024 · 被引用 124 次
- Vision-Language-Action Pretraining from Large-Scale Human VideosHao Luo, Yicheng Feng, Wanpeng Zhang, Sipeng Zheng 等ICML 2026 · 被引用 104 次
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 被引用 78 次
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 被引用 50 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
相关 Paper
- Modal-Enhanced Semantic Modeling for Fine-Grained 3D Human Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2024 · 被引用 3 次
- Hierarchical Semantics Alignment for 3D Human Motion RetrievalYang Yang, Haoyu Shi, Huaiwen ZhangSIGIR 2024 · 被引用 4 次
- Multi-Instance Multi-Label Learning for Text-motion RetrievalYang Yang, Liyuan Cao, Haoyu Shi, Huaiwen ZhangACM MM 2024 · 被引用 6 次
- Event-T2M: Event-level Conditioning for Complex Text-to-Motion SynthesisSeong-Eun Hong, JaeYoung Seon, Juyeong Hwang, JongHwan Shin 等ICLR 2026
- Sequence-Event Semantic Consistent Learning for Text-to-Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2025
