HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
Chengwen Zhang, Chun Yu, Borong Zhuang, Haopeng Jin, Qingyang Wan, Zhuojun Li, Zhe He, Zhoutong Ye, Yu Mei, Chang Liu, Weinan Shi, Yuanchun Shi
Abstract
Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) — determining who issued a command — is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34 m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on26
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 997 citations
- ABD-Net: Attentive but Diverse Person Re-IdentificationTianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan et al.ICCV 2019 · 544 citations
- Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal CorrespondencesHyunjong Park, Sanghoon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 248 citations
- Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic InterfacesRyo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati et al.CHI 2022 · 243 citations
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao et al.ICCV 2019 · 241 citations
Related papers
- I'M HOI: Inertia-Aware Monocular Capture of 3D Human-Object InteractionsChengfeng Zhao, Juze Zhang, Jiashen Du, Ziwei Shan et al.CVPR 2024 · 9 citations
- IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial FusionLizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun et al.CVPR 2026 · 1 citation
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
- Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range ConditionsZhiyang Lu, Wen Jiang, Tianren Wu, Zhichao Wang et al.AAAI 2026
- LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR SensorsYiming Ren, Chengfeng Zhao, Yannan He, Peishan Cong et al.IEEE VR 2023 · 50 citations
