XRF V2: A Dataset for Action Summarization with Wi-Fi Signals, and IMUs in Phones, Watches, Earbuds, and Glasses
Bo Lan, Pei Li, Jiaxi Yin, Yunpeng Song, Ge Wang, Han Ding, Jinsong Han, Fei Wang
Abstract
Human Action Recognition (HAR) plays a crucial role in applications such as health monitoring, smart home automation, and human-computer interaction. While HAR has been extensively studied, action summarization using Wi-Fi and IMU signals in smart-home environments , which involves identifying and summarizing continuous actions, remains an emerging task. This paper introduces the novel XRF V2 dataset, designed for indoor daily activity Temporal Action Localization (TAL) and action summarization. XRF V2 integrates multimodal data from Wi-Fi signals, IMU sensors (smartphones, smartwatches, headphones, and smart glasses), and synchronized video recordings, offering a diverse collection of indoor activities from 16 volunteers across three distinct environments. To tackle TAL and action summarization, we propose the XRFMamba neural network, which excels at capturing long-term dependencies in untrimmed sensory sequences and achieves the best performance with an average mAP of 78.74, outperforming the recent WiFiTAD by 5.49 points in mAP@avg while using 35% fewer parameters. In action summarization, we introduce a new metric, Response Meaning Consistency (RMC), to evaluate action summarization performance. And it achieves an average Response Meaning Consistency (mRMC) of 0.802. We envision XRF V2 as a valuable resource for advancing research in human action localization, action forecasting, pose estimation, multimodal foundation models pre-training, synthetic data generation, and more. The data and code are available at https://github.com/aiotgroup/XRFV2.
CCS Concepts: • Human-centered computing → Ubiquitous and mobile computing systems and tools.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 505eeb72-e0ec-4900-8358-508f813104f0Cited by top-tier papers2
- RF-HOI: Recognize Human-Object Interaction with Radio Frequency SignalsLihao Wang, Linlu Gao, Jiacan Yu, Yanyu Lin et al.UbiComp 2026
- Light but Sharp: SlimSTAD for Real-Time Action Detection from Sensor DataWei Cui, Lukai Fan, Zhenghua Chen, Min Wu et al.AAAI 2026
Builds on20
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- A decoder-only foundation model for time-series forecastingAbhimanyu Das, Weihao Kong, Rajat Sen, Yichen ZhouICML 2024 · 601 citations
- HACS: Human Action Clips and Segments Dataset for Recognition and Temporal LocalizationHang Zhao, Antonio Torralba, Lorenzo Torresani, Zhicheng YanICCV 2019 · 298 citations
- Towards 3D human pose construction using wifiWenjun Jiang, Hongfei Xue, Chenglin Miao, Shiyang Wang et al.MobiCom 2020 · 282 citations
Related papers
- XRF55: A Radio Frequency Dataset for Human Indoor Action AnalysisFei Wang, Yizhe Lv, Mengdie Zhu, Han Ding et al.UbiComp 2024 · 47 citations
- XGait: A Multi-Modality Wireless Sensing Dataset for Indoor Human Tracking and IdentificationWei Xu, Zhu Wang, Yifan Guo, Changlong Cheng et al.UbiComp 2026
- MoBind: Motion Binding for Fine-Grained IMU-Video Pose AlignmentDuc Duy Nguyen, Tat-Jun Chin, Minh HoaiCVPR 2026 · 1 citation
- RF-CM: Cross-Modal Framework for RF-enabled Few-Shot Human Activity RecognitionXuan Wang, Tong Liu, Chao Feng, Dingyi Fang et al.UbiComp 2023 · 18 citations
- UNI-FI: Integrated Multi-Task Wi-Fi SensingMengning Li, Wenye WangINFOCOM 2026 · 1 citation
