ASAP: Acoustic-Semantic Alignment with Prototypes for Open-world Activity Recognition and Understanding on Glasses
Changfei Dong, Qian Zhang, Dong Wang
摘要
Acoustic sensing-enabled smart glasses offer a privacy-preserving substrate for continuous activity understanding by probing near-body motion with inaudible acoustics. However, existing solutions are typically trained as closed-set classifiers: expanding the activity vocabulary requires costly data collection, and accuracy degrades sharply under cross-user and cross-environment shift. We present ASAP (Acoustic-Semantic Alignment with Prototypes), an open-world activity recognition system for ultrasonic smart glasses that targets practical near-neighbor vocabulary growth and robust generalization in daily life. ASAP converts ultrasonic echoes into motion-sensitive features and aligns them with text-derived activity prototypes in a shared acoustic-language embedding space. At inference time, ASAP integrates similarity-based rejection with two-stage seen/unseen routing, and supports rapid personalization via few-shot prototype fusion without retraining. Across 26 participants with 20 seen and 7 unseen activities, ASAP achieves 71.10% harmonic-mean accuracy under generalized zero-shot recognition, improves to 80.23% with few-shot fusion at k =3, and boosts cross-user generalization over a strong baseline by up to 18.8 points. For long-form continuous streams, ASAP further leverages an LLM as a post-hoc calibrator and journal generator, raising the harmonic mean to 80.87% (zero-shot) and 88.73% (few-shot) on in-the-wild long sessions, while user feedback shows that generated journals are preferred over raw recognition streams.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ActSonic: Recognizing Everyday Activities from Inaudible Acoustic Wave Around the BodySaif Mahmud, Vineet Parikh, Qikang Liang, Ke Li 等UbiComp 2025 · 被引用 24 次
- EchoLIFE: Zero-Shot In-Home ADL Recognition with LLM-Guided Active Acoustic SensingYubin Lan, Qian Zhang, Shukai Ma, Changfei Dong 等UbiComp 2026
- IMUZero: Zero-Shot Human Activity Recognition by Language-Based Cross Modality FusionJie Su, Fengtong Ge, Zhenyu Wen, Taotao Li 等UbiComp 2026 · 被引用 2 次
- A CNN-based Human Activity Recognition System Combining a Laser Feedback Interferometry Eye Movement Sensor and an IMU for Context-aware Smart GlassesJohannes Meyer, Adrian Frank, Thomas Schlebusch, Enkelejda KasneciUbiComp 2022 · 被引用 31 次
- PoseSonic: 3D Upper Body Pose Estimation Through Egocentric Acoustic Sensing on SmartglassesSaif Mahmud, Ke Li, Guilin Hu, Hao Chen 等UbiComp 2023 · 被引用 31 次
