MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
Duc Duy Nguyen, Tat-Jun Chin, Minh Hoai
Abstract
We aim to learn a joint representation between inertial measurement unit (IMU) signals and 2D pose sequences extracted from video, enabling accurate cross-modal retrieval, temporal synchronization, subject and body-part localization, and action recognition. To this end, we introduce MoBind, a hierarchical contrastive learning framework designed to address three challenges: (1) filtering out irrelevant visual background, (2) modeling structured multi-sensor IMU configurations, and (3) achieving fine-grained, sub-second temporal alignment. To isolate motion-relevant cues, MoBind aligns IMU signals with skeletal motion sequences rather than raw pixels. We further decompose full-body motion into local body-part trajectories, pairing each with its corresponding IMU to enable semantically grounded multi-sensor alignment. To capture detailed temporal correspondence, MoBind employs a hierarchical contrastive strategy that first aligns token-level temporal segments, then fuses local (body-part) alignment with global (body-wide) motion aggregation. Evaluated on mRi, TotalCapture, and EgoHumans, MoBind consistently outperforms strong baselines across all four tasks, demonstrating robust fine-grained temporal alignment while preserving coarse semantic consistency across modalities. Code is available at https://github.com/bbvisual/ MoBind.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f5cbab87-eaee-410f-841a-c7198d1e7bb7Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionTianyu Guo, Hong Liu, Zhan Chen, Mengyuan Liu et al.AAAI 2022 · 206 citations
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 158 citations
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian et al.ACM MM 2021 · 154 citations
- Contrastive Predictive Coding for Human Activity RecognitionHarish Haresamudram, Irfan A. Essa, Thomas PlötzUbiComp 2021 · 149 citations
Related papers
- Spatial-Related Sensors Matters: 3D Human Motion Reconstruction Assisted with Textual SemanticsXueyuan Yang, Chao Yao, Xiaojuan BanAAAI 2024 · 4 citations
- Vistar: Enhancing the Perception Capability of LLMs under Imprecise IMU-Text AlignmentYatong Chen, Chenzhi Hu, Bowen He, Ruijie Wang et al.KDD 2026
- Egocentric Action-Aware Inertial Localization in Point Clouds with Vision-Language GuidanceMingfang Zhang, Ryo Yonetani, Yifei Huang, Liangyang Ouyang et al.ICCV 2025
- IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial FusionLizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun et al.CVPR 2026 · 1 citation
- Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation LearningMinghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu et al.ACM MM 2023 · 6 citations
