Sequence-Event Semantic Consistent Learning for Text-to-Motion Retrieval
Haoyu Shi, Huaiwen Zhang
Abstract
Text-to-Motion Retrieval (TMR) is a challenging task to retrieve relevant motion sequences with the natural language description. Existing TMR methods primarily utilize single embeddings to represent and align text and motion sequences. However, real-world motion sequences typically contain multiple sequential actions with intricate semantics, which are hard to precisely capture by single embedding. Additionally, relying solely on naive contrastive training to capture high-level semantics may struggle to perceive and capture fine-grained action details necessary for precise text-motion alignment. In this work, we propose a novel Sequence-Event Semantic Consistent Learning (SECL) framework for 3D human motion retrieval. Specifically, we introduce a self-supervised learning strategy to incorporate fine-grained action details into the motion representations via the generative feedback from the diffusion model. We design a parameter-free sequence-level interaction to explore coarse-grained alignment and an event-level interaction that utilizes several learnable queries to capture event semantics in a shared learning manner for fine-grained alignment. Furthermore, an inter-consistency loss is introduced to align the event semantics between the motion and corresponding text, and an intra-diversity loss is designed to encourage event features to attend to different contents, effectively capturing the rich action information. Finally, we modify the traditional contrastive alignment objective and propose an importance-sampling strategy to emphasize harder negatives for discriminative representation learning. Extensive experiments show that our method significantly outperforms existing methods in text-to-motion retrieval and other challenging tasks, e.g., human interaction recognition and motion temporal localization.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Hierarchical Semantics Alignment for 3D Human Motion RetrievalYang Yang, Haoyu Shi, Huaiwen ZhangSIGIR 2024 · 4 citations
- Modal-Enhanced Semantic Modeling for Fine-Grained 3D Human Motion RetrievalHaoyu Shi, Huaiwen ZhangACM MM 2024 · 3 citations
- Multi-Instance Multi-Label Learning for Text-motion RetrievalYang Yang, Liyuan Cao, Haoyu Shi, Huaiwen ZhangACM MM 2024 · 6 citations
- Frame-wise Action Representations for Long Videos via Sequence Contrastive LearningMinghao Chen, Fangyun Wei, Chong Li, Deng CaiCVPR 2022 · 34 citations
- Video Moment Retrieval with Hierarchical Contrastive LearningBolin Zhang, Chao Yang, Bin Jiang, Xiaokang ZhouACM MM 2022 · 21 citations
