Multimodal Daily-Life Logging in Free-living Environment Using Non-Visual Egocentric Sensors on a Smartphone
Ke Sun, Chunyu Xia, Xinyu Zhang, Hao Chen, Charlie Jianzhong Zhang
摘要
Egocentric non-intrusive sensing of human activities of daily living (ADL) in free-living environments represents a holy grail in ubiquitous computing. Existing approaches, such as egocentric vision and wearable motion sensors, either can be intrusive or have limitations in capturing non-ambulatory actions. To address these challenges, we propose EgoADL, the first egocentric ADL sensing system that uses an in-pocket smartphone as a multi-modal sensor hub to capture body motion, interactions with the physical environment and daily objects using non-visual sensors (audio, wireless sensing, and motion sensors). We collected a 120-hour multimodal dataset and annotated 20-hour data into 221 ADL, 70 object interactions, and 91 actions. EgoADL proposes multi-modal frame-wise slow-fast encoders to learn the feature representation of multi-sensory data that characterizes the complementary advantages of different modalities and adapt a transformer-based sequence-to-sequence model to decode the time-series sensor signals into a sequence of words that represent ADL. In addition, we introduce a self-supervised learning framework that extracts intrinsic supervisory signals from the multi-modal sensing data to overcome the lack of labeling data and achieve better generalization and extensibility. Our experiments in free-living environments demonstrate that EgoADL can achieve comparable performance with video-based approaches, bringing the vision of ambient intelligence closer to reality.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- DiversityOne: A Multi-Country Smartphone Sensor Dataset for Everyday Life Behavior ModelingMatteo Busso, Andrea Bontempelli, Leonardo Javier Malcotti, Lakmal Meegahapola 等UbiComp 2025 · 被引用 12 次
- XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and GlassesBo Lan, Pei Li, Jiaxi Yin, Yunpeng Song 等UbiComp 2025 · 被引用 6 次
- Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture UnderstandingZhuoming Li, Aitong Liu, Mengxi Jia, Yubo Lu 等UbiComp 2026 · 被引用 1 次
- RF-HOI: Recognize Human-Object Interaction with Radio Frequency SignalsLihao Wang, Linlu Gao, Jiacan Yu, Yanyu Lin 等UbiComp 2026
相关 Paper
- Learning State-Aware Visual Representations from Audible InteractionsHimangi Mittal, Pedro Morgado, Unnat Jain, Abhinav GuptaNeurIPS 2022 · 被引用 30 次
- Leveraging Sound and Wrist Motion to Detect Activities of Daily Living with Commodity SmartwatchesSarnab Bhattacharya, Rebecca Adaimi, Edison ThomazUbiComp 2022 · 被引用 41 次
- HMotionGPT: Aligning Hand Motions and Natural Language for Activity Understanding with Smart RingsYang Gao, Dong She, Wolin Liang, Chiyue Wang 等UbiComp 2026
- Reading Recognition in the WildCharig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx 等NeurIPS 2025 · 被引用 9 次
- ActivitySeeker: Towards Collaborative Personalized Human Activity Discovery and Recognition on SmartphonesZhoutong Ye, Yanwen Huang, Chun Yu, Yuntao Wang 等CHI 2026 · 被引用 1 次
