MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
Duc Duy Nguyen, Tat-Jun Chin, Minh Hoai
摘要
We aim to learn a joint representation between inertial measurement unit (IMU) signals and 2D pose sequences extracted from video, enabling accurate cross-modal retrieval, temporal synchronization, subject and body-part localization, and action recognition. To this end, we introduce MoBind, a hierarchical contrastive learning framework designed to address three challenges: (1) filtering out irrelevant visual background, (2) modeling structured multi-sensor IMU configurations, and (3) achieving fine-grained, sub-second temporal alignment. To isolate motion-relevant cues, MoBind aligns IMU signals with skeletal motion sequences rather than raw pixels. We further decompose full-body motion into local body-part trajectories, pairing each with its corresponding IMU to enable semantically grounded multi-sensor alignment. To capture detailed temporal correspondence, MoBind employs a hierarchical contrastive strategy that first aligns token-level temporal segments, then fuses local (body-part) alignment with global (body-wide) motion aggregation. Evaluated on mRi, TotalCapture, and EgoHumans, MoBind consistently outperforms strong baselines across all four tasks, demonstrating robust fine-grained temporal alignment while preserving coarse semantic consistency across modalities. Code is available at https://github.com/bbvisual/ MoBind.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Contrastive Learning from Extremely Augmented Skeleton Sequences for Self-Supervised Action RecognitionTianyu Guo, Hong Liu, Zhan Chen, Mengyuan Liu 等AAAI 2022 · 被引用 206 次
- Skeleton-Contrastive 3D Action Representation LearningFida Mohammad Thoker, Hazel Doughty, Cees G. M. SnoekACM MM 2021 · 被引用 158 次
- Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionRuijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 等ACM MM 2021 · 被引用 154 次
- Contrastive Predictive Coding for Human Activity RecognitionHarish Haresamudram, Irfan A. Essa, Thomas PlötzUbiComp 2021 · 被引用 149 次
相关 Paper
- Spatial-Related Sensors Matters: 3D Human Motion Reconstruction Assisted with Textual SemanticsXueyuan Yang, Chao Yao, Xiaojuan BanAAAI 2024 · 被引用 4 次
- Vistar: Enhancing the Perception Capability of LLMs under Imprecise IMU-Text AlignmentYatong Chen, Chenzhi Hu, Bowen He, Ruijie Wang 等KDD 2026
- Egocentric Action-Aware Inertial Localization in Point Clouds with Vision-Language GuidanceMingfang Zhang, Ryo Yonetani, Yifei Huang, Liangyang Ouyang 等ICCV 2025
- IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial FusionLizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun 等CVPR 2026 · 被引用 1 次
- Fine-Grained Spatiotemporal Motion Alignment for Contrastive Video Representation LearningMinghao Zhu, Xiao Lin, Ronghao Dang, Chengju Liu 等ACM MM 2023 · 被引用 6 次
