Transformer with Controlled Attention for Synchronous Motion Captioning
Karim Radouane, Sylvie Ranwez, Julien Lagarde, Andon Tchechmedjiev
Abstract
In this paper, we address a challenging task, synchronous motion captioning, that aim to generate a language description synchronized with human motion sequences. This task pertains to numerous applications, such as aligned sign language transcription and unsupervised action segmentation and temporal grounding. Our method introduces mechanisms to control self- and cross-attention distributions of the Transformer, allowing interpretability and aligned text generation. We achieve this through masking strategies and structuring losses that push the model to maximize attention only on the most important frames contributing to the generation of a motion word. These constraints aim to prevent undesired mixing of information in attention maps and to provide a monotonic attention distribution across tokens. Thus, the cross attentions of tokens are used for progressive text generation in synchronization with human motion sequences. We demonstrate the superior performance of our approach through evaluation on the two available benchmark datasets, KIT-ML and HumanML3D. As visual evaluation is essential for this task, we provide a comprehensive set of animated visual illustrations of the output of synchronous text generation in the code repository.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- MotionGPT: Human Motion as a Foreign LanguageBiao Jiang, Xin Chen, Wen Liu, Jingyi Yu et al.NeurIPS 2023 · 698 citations
- Synthesis of Compositional Animations from Textual DescriptionsAnindita Ghosh, Noshaba Cheema, Cennet Oguz, Christian Theobalt et al.ICCV 2021 · 226 citations
- TMR: Text-to-Motion Retrieval Using Contrastive 3D Human Motion SynthesisMathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 192 citations
- Aligning Subtitles in Sign Language VideosHannah Bull, Triantafyllos Afouras, Gül Varol, Samuel Albanie et al.ICCV 2021 · 39 citations
- Executing your Commands via Motion Diffusion in Latent SpaceXin Chen, Biao Jiang, Wen Liu, Zilong Huang et al.CVPR 2023
Related papers
- SnapMoGen: Human Motion Generation from Expressive TextsChuan Guo, Inwoo Hwang, Jian Wang, Bing ZhouNeurIPS 2025 · 50 citations
- Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic GraphsPeng Jin, Yang Wu, Yanbo Fan, Zhongqian Sun et al.NeurIPS 2023 · 57 citations
- HGM³: Hierarchical Generative Masked Motion Modeling with Hard Token MiningMinjae Jeong, Yechan Hwang, Jaejin Lee, Sungyoon Jung et al.ICLR 2025
- Motion-Aligned Word Embeddings for Text-to-Motion GenerationKe Han, Yueming Lyu, Nicu SebeICLR 2026
- Towards Unified Human Motion-Language Understanding via Sparse Interpretable CharacterizationGuangtao Lyu, Chenghao Xu, Jiexi Yan, Muli Yang et al.ICLR 2025
