Tri-Modal Motion Retrieval by Learning a Joint Embedding Space
Kangning Yin, Shihao Zou, Yuxuan Ge, Zheng Tian
Abstract
Information retrieval is an ever-evolving and crucial re-search domain. The substantial demand for high-quality human motion data especially in online acquirement has led to a surge in human motion research works. Prior works have mainly concentrated on dual-modality learning, such as text and motion tasks, but three-modality learning has been rarely explored. Intuitively, an extra introduced modality can enrich a model's application scenario, and more importantly, an adequate choice of the extra modality can also act as an intermediary and enhance the alignment between the other two disparate modalities. In this work, we introduce LAVIMO (LAnguage-VIdeo-MOtion alignment), a novel framework for three-modality learning integrating human-centric videos as an additional modality, thereby ef-fectively bridging the gap between text and motion. More-over, our approach leverages a specially designed attention mechanism to foster enhanced alignment and synergistic effects among text, video, and motion modalities. Empirically, our results on the HumanML3D and KIT-ML datasets show that LAVIMO achieves state-of-the-art performance in various motion-related cross-modal retrieval tasks, in-cluding text-to-motion, motion-to-text, video-to-motion and motion-to-video. Our project webpage can be found in https://lavimo2023.github.io/LAVIMO/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6926a052-6094-4272-a488-2f0a26903e54Cited by top-tier papers7
- MotionGPT3: Human Motion as a Second ModalityBingfan Zhu, Biao Jiang, Sunyi Wang, Shixiang Tang et al.ICLR 2026 · 43 citations
- SGAR: Structural Generative Augmentation for 3D Human Motion RetrievalJiahang Zhang, Lilang Lin, Shuai Yang, Jiaying LiuNeurIPS 2025 · 7 citations
- MMGeo: Multimodal Compositional Geo-Localization for UAVsYuxiang Ji, Boyong He, Zhuoyue Tan, Liaoni WuICCV 2025 · 5 citations
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingThomas Kreutz, Max Mühlhäuser, Alejandro Sánchez GuineaICCV 2025 · 1 citation
- MonSTeR: A Unified Model for Motion, Scene, Text RetrievalLuca Collorone, Matteo Gioia, Massimiliano Pappa, Paolo Leoni et al.ICCV 2025 · 1 citation
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
Related papers
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language ModelsHaidong Xu, Guangwei Xu, Zhedong Zheng, Xiatian Zhu et al.NeurIPS 2025 · 5 citations
- Hierarchical Semantics Alignment for 3D Human Motion RetrievalYang Yang, Haoyu Shi, Huaiwen ZhangSIGIR 2024 · 4 citations
- Multi-Instance Multi-Label Learning for Text-motion RetrievalYang Yang, Liyuan Cao, Haoyu Shi, Huaiwen ZhangACM MM 2024 · 6 citations
- MotionBind: Multi-Modal Human Motion Alignment for Retrieval, Recognition, and GenerationKaleab Alemayehu Kinfu, René VidalNeurIPS 2025 · 4 citations
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 192 citations
