VAX: Using Existing Video and Audio-based Activity Recognition Models to Bootstrap Privacy-Sensitive Sensors
Prasoon Patidar, Mayank Goel, Yuvraj Agarwal
Abstract
The use of audio and video modalities for Human Activity Recognition (HAR) is common, given the richness of the data and the availability of pre-trained ML models using a large corpus of labeled training data. However, audio and video sensors also lead to significant consumer privacy concerns. Researchers have thus explored alternate modalities that are less privacy-invasive such as mmWave doppler radars, IMUs, motion sensors. However, the key limitation of these approaches is that most of them do not readily generalize across environments and require significant in-situ training data. Recent work has proposed cross-modality transfer learning approaches to alleviate the lack of trained labeled data with some success. In this paper, we generalize this concept to create a novel system called VAX (Video/Audio to 'X'), where training labels acquired from existing Video/Audio ML models are used to train ML models for a wide range of 'X' privacy-sensitive sensors. Notably, in VAX, once the ML models for the privacy-sensitive sensors are trained, with little to no user involvement, the Audio/Video sensors can be removed altogether to protect the user's privacy better. We built and deployed VAX in ten participants' homes while they performed 17 common activities of daily living. Our evaluation results show that after training, VAX can use its onboard camera and microphone to detect approximately 15 out of 17 activities with an average accuracy of 90%. For these activities that can be detected using a camera and a microphone, VAX trains a per-home model for the privacy-preserving sensors. These models (average accuracy = 84%) require no in-situ user input. In addition, when VAX is augmented with just one labeled instance for the activities not detected by the VAX A/V pipeline ( 2 out of 17), it can detect all 17 activities with an average accuracy of 84%. Our results show that VAX is significantly better than a baseline supervised-learning approach of using one labeled instance per activity in each home (average accuracy of 79%) since VAX reduces the user burden of providing activity labels by 8x ( 2 labels vs. 17 labels).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 232d0897-acf3-4ff8-bf1a-fe27a0bac568Cited by top-tier papers5
- PrISM-Observer: Intervention Agent to Help Users Perform Everyday Procedures Sensed using a SmartwatchRiku Arakawa, Hiromu Yakura, Mayank GoelUIST 2024 · 20 citations
- Poison to Cure: Privacy-preserving Wi-Fi Multi-User Sensing via Data PoisoningJingzhi Hu, Xin Li, Jin Gan, Jun LuoMobiCom 2025 · 4 citations
- Scaling Context-Aware Task Assistants that Learn from Demonstration and Adapt through Mixed-Initiative DialogueRiku Arakawa, Prasoon Patidar, Will Page, Jill Lehman et al.UIST 2025 · 3 citations
- After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants from Large-Scale Simulated Persona InteractionsZiyi Xuan, Yiwen Wu, Zhaoyang Yan, Vinod Namboodiri et al.UbiComp 2026
- OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video AnalysisPrasoon Patidar, Riku Arakawa, Ricardo Graça, Ruben Moutinho et al.UbiComp 2026
Builds on21
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 2,927 citations
- TSM: Temporal Shift Module for Efficient Video UnderstandingJi Lin, Chuang Gan, Song HanICCV 2019 · 2,049 citations
- Revisiting Skeleton-based Action RecognitionHaodong Duan, Yue Zhao, Kai Chen, Dahua Lin et al.CVPR 2022 · 752 citations
- TAM: Temporal Adaptive Module for Video RecognitionZhaoyang Liu, Limin Wang, Wayne Wu, Chen Qian et al.ICCV 2021 · 356 citations
Related papers
- Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity RecognitionKaran Ahuja, Yue Jiang, Mayank Goel, Chris HarrisonCHI 2021 · 118 citations
- mmHolmes: Amodal Millimeter-wave Sensing by Understanding Human KineticsKun Liang, Yaxuan Li, He Hao, Hongli Zeng et al.UbiComp 2025 · 2 citations
- IMU2Doppler: Cross-Modal Domain Adaptation for Doppler-based Activity Recognition Using IMU DataSejal Bhalla, Mayank Goel, Rushil KhuranaUbiComp 2022 · 42 citations
- Midas: Generating mmWave Radar Data from Videos for Training Pervasive and Privacy-preserving Human Sensing TasksKaikai Deng, Dong Zhao, Qiaoyue Han, Zihan Zhang et al.UbiComp 2023 · 36 citations
- Adapting Pretrained Large Vision Models for Sensor-based Activity RecognitionYize Cai, Rui Feng, Kunlin Cai, Yunhuai Liu et al.UbiComp 2026
