OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis
Prasoon Patidar, Riku Arakawa, Ricardo Graça, Ruben Moutinho, Adriano Soares, Ana Vasconcelos, Filippo Talami, Joana Couto da Silva, Inês Silva, Cristina Mendes-Santos, Mayank Goel, Yuvraj Agarwal
Abstract
Deploying human activity recognition (HAR) at home is still rare because sensor signals vary wildly across houses, people, and time, essentially requiring in-situ data collection and training. Prior approaches use cameras to generate training labels for privacy-preserving sensors (LiDAR, RADAR, Thermal), but this forces sensors to detect predefined activities that cameras can see yet the sensors themselves cannot reliably distinguish. In this work, we introduce OrganicHAR, an activity discovery framework that inverts this relationship by placing sensor capabilities at the center of activity discovery. Our approach identifies naturally occurring signal patterns using privacy-preserving sensors, leverages Vision Language Models (VLMs) only during these key moments for scene understanding, and discovers discrete activity labels at granularities that these sensors can reliably detect. Our evaluation with 12 participants demonstrates OrganicHAR's effectiveness: it achieves 79% accuracy for coarse (4-5) activities using only basic ambient sensors (radar, lidar, thermal arrays), and 73% accuracy for fine-grained (8-9) activities when a wearable IMU, depth, and pose sensor are added. OrganicHAR maintains 77% accuracy on average across configurations while discovering 4-8 categories per user (15 across all users) tailored to each environment and sensor capabilities. By triggering video processing only at key moments identified by local sensors, we reduce queries to VLM by 90%, enabling practical and privacy-preserving activity recognition in natural settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 153b661a-cbb3-46f3-b2cd-dbfbb2a9c896Builds on22
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez et al.ICLR 2024 · 154 citations
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity RecognitionHyeokHyen Kwon, Catherine Tong, Harish Haresamudram, Yan Gao et al.UbiComp 2020 · 153 citations
- Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity RecognitionKaran Ahuja, Yue Jiang, Mayank Goel, Chris HarrisonCHI 2021 · 118 citations
- ColloSSL: Collaborative Self-Supervised Learning for Human Activity RecognitionYash Jain, Chi Ian Tang, Chulhong Min, Fahim Kawsar et al.UbiComp 2022 · 113 citations
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison et al.CHI 2023 · 103 citations
Related papers
- VAX: Using Existing Video and Audio-based Activity Recognition Models to Bootstrap Privacy-Sensitive SensorsPrasoon Patidar, Mayank Goel, Yuvraj AgarwalUbiComp 2023 · 12 citations
- RF-CM: Cross-Modal Framework for RF-enabled Few-Shot Human Activity RecognitionXuan Wang, Tong Liu, Chao Feng, Dingyi Fang et al.UbiComp 2023 · 18 citations
- HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and ReasoningChuhao Zhou, Jianfei YangNeurIPS 2025 · 4 citations
- mmHolmes: Amodal Millimeter-wave Sensing by Understanding Human KineticsKun Liang, Yaxuan Li, He Hao, Hongli Zeng et al.UbiComp 2025 · 2 citations
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingThomas Kreutz, Max Mühlhäuser, Alejandro Sánchez GuineaICCV 2025 · 1 citation
