COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
Baiyu Chen, Wilson Wongso, Zechen Li, Yonchanok Khaokaew, Hao Xue, Flora D. Salim
摘要
The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocentric video-based models capture rich semantic information and have demonstrated strong performance in human activity recognition (HAR), but their high power consumption, privacy concerns, and dependence on lighting limit their feasibility for continuous on-device recognition. In contrast, inertial measurement unit (IMU) sensors offer an energy-efficient, privacy-preserving alternative, yet lack large-scale annotated datasets, leading to weaker generalization. To bridge this gap, we propose COMODO, a cross-modal self-supervised distillation framework that transfers semantic knowledge from video to IMU without requiring labels. COMODO leverages a pretrained and frozen video encoder to construct a dynamic instance queue to align the feature distributions of video and IMU embeddings. This enables the IMU encoder to inherit rich semantic structure from video while maintaining its efficiency for real-world applications. Experiments on multiple egocentric HAR datasets show that COMODO consistently improves downstream performance, matching or surpassing fully supervised models, and demonstrating strong cross-dataset generalization. Benefiting from its simplicity and flexibility, COMODO is compatible with diverse pretrained video and time-series models, offering the potential to leverage more powerful teacher and student foundation models in future ubiquitous computing research. The code is available at this repository: https://github.com/cruiseresearchgroup/COMODO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CauKer: Classification Time Series Foundation Models Can Be Pretrained on Synthetic DataShifeng Xie, Vasilii Feofanov, Jianfeng Zhang, Themis Palpanas 等ICLR 2026 · 被引用 15 次
- MoPFormer: Motion-Primitive Transformer for Wearable-Sensor Activity RecognitionHao Zhang, Zhan Zhuang, Xuehao Wang, Xiaodong Yang 等NeurIPS 2025 · 被引用 11 次
- DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged LearningJunho Yoon, Jaemo Jeong, Hyunju Kim, Dongman LeeCVPR 2026
- ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM AgentsZechen Li, Baiyu Chen, Hao Xue, Flora D. SalimACL 2026
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series ForecastingHaoyi Zhou, Shanghang Zhang, Jieqi Peng, Shuai Zhang 等AAAI 2021 · 被引用 7,289 次
- Are Transformers Effective for Time Series Forecasting?Ailing Zeng, Muxi Chen, Lei Zhang, Qiang XuAAAI 2023 · 被引用 3,619 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
相关 Paper
- EgoDistill: Egocentric Head Motion Distillation for Efficient Video UnderstandingShuhan Tan, Tushar Nagarajan, Kristen GraumanNeurIPS 2023 · 被引用 44 次
- Adapting Pretrained Large Vision Models for Sensor-based Activity RecognitionYize Cai, Rui Feng, Kunlin Cai, Yunhuai Liu 等UbiComp 2026
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 被引用 1 次
- ActivitySeeker: Towards Collaborative Personalized Human Activity Discovery and Recognition on SmartphonesZhoutong Ye, Yanwen Huang, Chun Yu, Yuntao Wang 等CHI 2026 · 被引用 1 次
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal AttentionKatsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro OkadaACM MM 2021 · 被引用 9 次
