EMBridge: Enhancing Gesture Generalization from EMG Signals Through Cross-modal Representation Learning
Wenhui Cui, Christopher Michael Sandino, Hadi Pouransari, Ran Liu, Juri Minxha, Ellen L. Zippi, Erdrin Azemi, Behrooz Mahasseni
摘要
Hand gesture classification using high-quality structured data such as videos, images, and hand skeletons is a well-explored problem in computer vision. Alternatively, leveraging low-power, cost-effective bio-signals, e.g. surface electromyography (sEMG), allows for continuous gesture prediction on wearable devices. In this work, we aim to enhance EMG representation quality by aligning it with embeddings obtained from structured, high-quality modalities that provide richer semantic guidance, ultimately enabling zero-shot gesture generalization. Specifically, we propose EMBridge, a cross-modal representation learning framework that bridges the modality gap between EMG and pose. EMBridge learns high-quality EMG representations by introducing a Querying Transformer (Q-Former), a masked pose reconstruction loss, and a community-aware soft contrastive learning objective that aligns the relative geometry of the embedding spaces. We evaluate EMBridge on both in-distribution and unseen gesture classification tasks and demonstrate consistent performance gains over all baselines. To the best of our knowledge, EMBridge is the first cross-modal representation learning framework to achieve zero-shot gesture classification from wearable EMG signals, showing potential toward real-world gesture recognition on wearable devices.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- Reconstructing Hands in 3D with TransformersGeorgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa 等CVPR 2024 · 被引用 110 次
相关 Paper
- Reading Your Actions: Learning Generalizable Action Representations via Pre-training AEMGZhenghao Huang, Huilin Yao, Kaikai Wang, Lin ShuCVPR 2026
- New Synthetic Goldmine: Hand Joint Angle-Driven EMG Data Generation Framework for Micro-Gesture RecognitionNana Wang, Suli Wang, Gen Li, Pengfei Ren 等AAAI 2026
- EmergentBridge: Improving Zero-Shot Cross-Modal Transfer in Unified Multimodal Embedding ModelsJincheng Xie, Xingchen Xiao, Runheng Liu, Zhongyi Huang 等KDD 2026 · 被引用 1 次
- Learning Unseen Emotions from Gestures via Semantically-Conditioned Zero-Shot Perception with Adversarial AutoencodersAbhishek Banerjee, Uttaran Bhattacharya, Aniket BeraAAAI 2022 · 被引用 19 次
- SpGesture: Source-Free Domain-adaptive sEMG-based Gesture Recognition with Jaccard Attentive Spiking Neural NetworkWeiyu Guo, Ying Sun, Yijie Xu, Ziyue Qiao 等NeurIPS 2024 · 被引用 17 次
