SensorLM: Learning the Language of Wearable Sensors
Yuwei Zhang, Kumar Ayush, Siyuan Qiao, A. Ali Heydari, Girish Narayanswamy, Max Xu, Ahmed Metwally, Jinhua Xu, Jake Garrison, Xuhai Orson Xu, Tim Althoff, Yun Liu
摘要
We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and interpreting sensor data with language remains challenging due to the lack of paired, richly annotated sensor-text descriptions in uncurated, real-world wearable data. We introduce a hierarchical caption generation pipeline designed to capture statistical, structural, and semantic information from sensor data. This approach enabled the curation of the largest sensor-language dataset to date, comprising over 59.7 million hours of data from more than 103,000 people. Furthermore, SensorLM extends prominent multimodal pretraining architectures (e.g., CLIP, CoCa) and recovers them as specific variants within a generic architecture. Extensive experiments on real-world tasks in human activity analysis and healthcare verify the superior performance of SensorLM over state-of-the-art in zero-shot recognition, few-shot learning, and cross-modal retrieval. SensorLM also demonstrates intriguing capabilities including scaling behaviors, label efficiency, sensor captioning, and zero-shot generalization to unseen tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Bloom: Designing for LLM-Augmented Behavior Change InteractionsMatthew Jörke, Defne Genç, Valentin Teutschbein, Shardul Sapkota 等CHI 2026 · 被引用 6 次
- MIND: Empowering Mental Health Clinicians with Multimodal Data Insights through a Narrative DashboardRuishi Zou, Shiyu Xu, Margaret E. Morris, Jihan Ryu 等CHI 2026 · 被引用 2 次
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And OutlookSizhen Bian, Mengxi Liu, Lala Shakti Swarup Ray, Bo Zhou 等UbiComp 2026 · 被引用 2 次
它引用的顶会 Paper16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- SimVLM: Simple Visual Language Model Pretraining with Weak SupervisionZirui Wang, Jiahui Yu, Adams Wei Yu, Zihang Dai 等ICLR 2022 · 被引用 950 次
相关 Paper
- SleepLM: Natural-Language Intelligence for Human SleepZongzhe Xu, Zitao Shuai, Eideen Mozaffari, Ravi Aysola 等ICML 2026 · 被引用 10 次
- SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity RecognitionZechen Li, Shohreh Deldari, Linyao Chen, Hao Xue 等EMNLP 2025 · 被引用 9 次
- Scaling Wearable Foundation ModelsGirish Narayanswamy, Xin Liu, Kumar Ayush, Yuzhe Yang 等ICLR 2025
- CLEP: Contrastive Language-Pose PretrainingSen Jia, Huayu Wang, Hsiang-Wei Huang, Zhaochong An 等CVPR 2026
- GOAT: A Generalized Cross-Dataset Activity Recognition Framework with Natural Language SupervisionShenghuan Miao, Ling ChenUbiComp 2025 · 被引用 13 次
