HiSync: Spatio-Temporally Aligning Hand Motion from Wearable IMU and On-Robot Camera for Command Source Identification in Long-Range HRI
Chengwen Zhang, Chun Yu, Borong Zhuang, Haopeng Jin, Qingyang Wan, Zhuojun Li, Zhe He, Zhoutong Ye, Yu Mei, Chang Liu, Weinan Shi, Yuanchun Shi
摘要
Long-range Human-Robot Interaction (HRI) remains underexplored. Within it, Command Source Identification (CSI) — determining who issued a command — is especially challenging due to multi-user and distance-induced sensor ambiguity. We introduce HiSync, an optical-inertial fusion framework that treats hand motion as binding cues by aligning robot-mounted camera optical flow with hand-worn IMU signals. We first elicit a user-defined (N=12) gesture set and collect a multimodal command gesture dataset (N=38) in long-range multi-user HRI scenarios. Next, HiSync extracts frequency-domain hand motion features from both camera and IMU data, and a learned CSINet denoises IMU readings, temporally aligns modalities, and performs distance-aware multi-window fusion to compute cross-modal similarity of subtle, natural gestures, enabling robust CSI. In three-person scenes up to 34 m, HiSync achieves 92.32% CSI accuracy, outperforming the prior SOTA by 48.44%. HiSync is also validated on real-robot deployment. By making CSI reliable and natural, HiSync provides a practical primitive and design guidance for public-space HRI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper26
- Omni-Scale Feature Learning for Person Re-IdentificationKaiyang Zhou, Yongxin Yang, Andrea Cavallaro, Tao XiangICCV 2019 · 被引用 997 次
- ABD-Net: Attentive but Diverse Person Re-IdentificationTianlong Chen, Shaojin Ding, Jingyi Xie, Ye Yuan 等ICCV 2019 · 被引用 544 次
- Learning by Aligning: Visible-Infrared Person Re-identification using Cross-Modal CorrespondencesHyunjong Park, Sanghoon Lee, Junghyup Lee, Bumsub HamICCV 2021 · 被引用 248 次
- Augmented Reality and Robotics: A Survey and Taxonomy for AR-enhanced Human-Robot Interaction and Robotic InterfacesRyo Suzuki, Adnan Karim, Tian Xia, Hooman Hedayati 等CHI 2022 · 被引用 243 次
- Global-Local Temporal Representations for Video Person Re-IdentificationJianing Li, Shiliang Zhang, Jingdong Wang, Wen Gao 等ICCV 2019 · 被引用 241 次
相关 Paper
- I'M HOI: Inertia-Aware Monocular Capture of 3D Human-Object InteractionsChengfeng Zhao, Juze Zhang, Jiashen Du, Ziwei Shan 等CVPR 2024 · 被引用 9 次
- IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial FusionLizhou Lin, Songpengcheng Xia, Zengyuan Lai, Lan Sun 等CVPR 2026 · 被引用 1 次
- How Humans Naturally Refer to Targets: Understanding Multimodal Instruction Patterns in Human-Robot InteractionLesong Jia, Makayla Chang, Yu Liu, Na DuCHI 2026
- Walking Further: Semantic-Aware Multimodal Gait Recognition Under Long-Range ConditionsZhiyang Lu, Wen Jiang, Tianren Wu, Zhichao Wang 等AAAI 2026
- LiDAR-aid Inertial Poser: Large-scale Human Motion Capture by Sparse Inertial and LiDAR SensorsYiming Ren, Chengfeng Zhao, Yannan He, Peishan Cong 等IEEE VR 2023 · 被引用 50 次
