OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis
Prasoon Patidar, Riku Arakawa, Ricardo Graça, Ruben Moutinho, Adriano Soares, Ana Vasconcelos, Filippo Talami, Joana Couto da Silva, Inês Silva, Cristina Mendes-Santos, Mayank Goel, Yuvraj Agarwal
摘要
Deploying human activity recognition (HAR) at home is still rare because sensor signals vary wildly across houses, people, and time, essentially requiring in-situ data collection and training. Prior approaches use cameras to generate training labels for privacy-preserving sensors (LiDAR, RADAR, Thermal), but this forces sensors to detect predefined activities that cameras can see yet the sensors themselves cannot reliably distinguish. In this work, we introduce OrganicHAR, an activity discovery framework that inverts this relationship by placing sensor capabilities at the center of activity discovery. Our approach identifies naturally occurring signal patterns using privacy-preserving sensors, leverages Vision Language Models (VLMs) only during these key moments for scene understanding, and discovers discrete activity labels at granularities that these sensors can reliably detect. Our evaluation with 12 participants demonstrates OrganicHAR's effectiveness: it achieves 79% accuracy for coarse (4-5) activities using only basic ambient sensors (radar, lidar, thermal arrays), and 73% accuracy for fine-grained (8-9) activities when a wearable IMU, depth, and pose sensor are added. OrganicHAR maintains 77% accuracy on average across configurations while discovering 4-8 categories per user (15 across all users) tailored to each environment and sensor capabilities. By triggering video processing only at key moments identified by local sensors, we reduce queries to VLM by 90%, enabling practical and privacy-preserving activity recognition in natural settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Vision-Language Models are Zero-Shot Reward Models for Reinforcement LearningJuan Rocamonde, Victoriano Montesinos, Elvis Nava, Ethan Perez 等ICLR 2024 · 被引用 154 次
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity RecognitionHyeokHyen Kwon, Catherine Tong, Harish Haresamudram, Yan Gao 等UbiComp 2020 · 被引用 153 次
- Vid2Doppler: Synthesizing Doppler Radar Data from Videos for Training Privacy-Preserving Activity RecognitionKaran Ahuja, Yue Jiang, Mayank Goel, Chris HarrisonCHI 2021 · 被引用 118 次
- ColloSSL: Collaborative Self-Supervised Learning for Human Activity RecognitionYash Jain, Chi Ian Tang, Chulhong Min, Fahim Kawsar 等UbiComp 2022 · 被引用 113 次
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and EarbudsVimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harrison 等CHI 2023 · 被引用 103 次
相关 Paper
- VAX: Using Existing Video and Audio-based Activity Recognition Models to Bootstrap Privacy-Sensitive SensorsPrasoon Patidar, Mayank Goel, Yuvraj AgarwalUbiComp 2023 · 被引用 12 次
- RF-CM: Cross-Modal Framework for RF-enabled Few-Shot Human Activity RecognitionXuan Wang, Tong Liu, Chao Feng, Dingyi Fang 等UbiComp 2023 · 被引用 18 次
- HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and ReasoningChuhao Zhou, Jianfei YangNeurIPS 2025 · 被引用 4 次
- mmHolmes: Amodal Millimeter-wave Sensing by Understanding Human KineticsKun Liang, Yaxuan Li, He Hao, Hongli Zeng 等UbiComp 2025 · 被引用 2 次
- DeSPITE: Exploring Contrastive Deep Skeleton-Pointcloud-IMU-Text Embeddings for Advanced Point Cloud Human Activity UnderstandingThomas Kreutz, Max Mühlhäuser, Alejandro Sánchez GuineaICCV 2025 · 被引用 1 次
