Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors
Wenqiang Chen, Jiaxuan Cheng, Leyao Wang, Wei Zhao, Wojciech Matusik
Abstract
Visual Question-Answering, a technology that generates textual responses from an image and natural language question, has progressed significantly. Notably, it can aid in tracking and inquiring about daily activities, crucial in healthcare monitoring, especially for elderly patients or those with memory disabilities. However, video poses privacy concerns and has a limited field of view. This paper presents Sensor2Text, a model proficient in tracking daily activities and engaging in conversations using wearable sensors. The approach outlined here tackles several challenges, including low information density in wearable sensor data, insufficiency of single wearable sensors in human activities recognition, and model's limited capacity for Question-Answering and interactive conversations. To resolve these obstacles, transfer learning and student-teacher networks are utilized to leverage knowledge from visual-language models. Additionally, an encoder-decoder neural network model is devised to jointly process language and sensor data for conversational purposes. Furthermore, Large Language Models are also utilized to enable interactive capabilities. The model showcases the ability to identify human activities and engage in Q&A dialogues using various wearable sensor modalities. It performs comparably to or better than existing visual-language models in both captioning and conversational tasks. To our knowledge, this represents the first model capable of conversing about wearable sensor data, offering an innovative approach to daily activity tracking that addresses privacy and field-of-view limitations associated with current vision-based solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a1b9d46a-1d6e-4c3e-adeb-04ac3b385cccCited by top-tier papers6
- SensorLM: Learning the Language of Wearable SensorsYuwei Zhang, Kumar Ayush, Siyuan Qiao, A. Ali Heydari et al.NeurIPS 2025 · 75 citations
- GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and WellbeingAkshat Choube, Ha Le, Jiachen Li, Kaixin Ji et al.UbiComp 2025 · 10 citations
- SensorChat: Answering Qualitative and Quantitative Questions during Long-term Multimodal Sensor InteractionsXiaofan Yu, Lanxiang Hu, Benjamin Z. Reichman, Dylan Chu et al.UbiComp 2025 · 4 citations
- RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural LanguageSubrata Biswas, Mohammad Nur Hossain Khan, Bashima IslamEMNLP 2025 · 3 citations
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou et al.UbiComp 2026 · 2 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
Related papers
- Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition - And Ways to Overcome ThemHarish Haresamudram, Apoorva Beedu, Mashfiqui Rabbi, Sankalita Saha et al.AAAI 2025 · 11 citations
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue et al.EMNLP 2025 · 9 citations
- Adapting Pretrained Large Vision Models for Sensor-based Activity RecognitionYize Cai, Rui Feng, Kunlin Cai, Yunhuai Liu et al.UbiComp 2026
- Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal AttentionKatsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro OkadaACM MM 2021 · 9 citations
- Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted InputsMinh Tran, Hongda Mao, Qingshuang Chen, Yelin KimICCV 2025 · 1 citation
