EchoScriptor: Automatic Lifelogging Narratives via Activity-Based Audio-Language Model
Kaylee Yaxuan Li, Xinghao Zhou, Haizhong Zheng, Kang G. Shin, Alanson P. Sample
Abstract
Automatic, camera-free lifelogging offers new opportunities for memory rehabilitation, personal informatics, and assistive technologies. However, most existing approaches limit daily activities to isolated event labels, offering little context and lacking the narrative coherence essential for effective lifelogging. Recent advances in audio–language models combine foundation audio processing with language-based reasoning, enabling open-ended sound understanding. We introduce EchoScriptor, an end-to-end system that transforms raw in-home audio into context-aware natural-language descriptions, generating coherent narrative lifelogs of activities and acoustic contexts. In moment-level evaluation, EchoScriptor achieved 94.15% activity recognition and 89.25% background recognition accuracy, and at the summary level, achieved an F1 score of 0.92, outperforming the classifier+LLM baseline. In our user study with 20 participants across 10 household activity videos, EchoScriptor summaries were consistently rated highly, approaching the perceived utility of human-written ones. By advancing from event detection to narrative understanding, EchoScriptor establishes a significant step toward automated, unobtrusive, context-aware lifelogging technologies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7f9bb914-1ecf-47ad-a4a3-e94b31082d43Related papers
- EchoLIFE: Zero-Shot In-Home ADL Recognition with LLM-Guided Active Acoustic SensingYubin Lan, Qian Zhang, Shukai Ma, Changfei Dong et al.UbiComp 2026
- SoundTrace: Integrating Temporal Context and Episodic Memory for Real-Time Sound RecognitionDhruv Jain, Jason MillerUbiComp 2026
- Automated Class Discovery and One-Shot Interactions for Acoustic Activity RecognitionJason Wu, Chris Harrison, Jeffrey P. Bigham, Gierad LaputCHI 2020 · 52 citations
- Ok Google, What Am I Doing?: Acoustic Activity Recognition Bounded by Conversational Assistant InteractionsRebecca Adaimi, Howard Yong, Edison ThomazUbiComp 2021 · 29 citations
- AutoLife: Automatic Life Journaling with Smartphones and LLMsHuatao Xu, Zilin Zeng, Panrong Tong, Mo Li et al.UbiComp 2026 · 6 citations
