OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question Answering
Jiahao Nick Li, Zhuohao Jerry Zhang, Jiaju Ma
Abstract
People often capture memories through photos, screenshots, and videos. While existing AI-based tools enable querying this data using natural language, they only support retrieving individual pieces of information like certain objects in photos, and struggle with answering more complex queries that involve interpreting interconnected memories like sequential events. We conducted a one-month diary study to collect realistic user queries and generated a taxonomy of necessary contextual information for integrating with captured memories. We then introduce OmniQuery, a novel system that is able to answer complex personal memory-related questions that require extracting and inferring contextual information. OmniQuery augments individual captured memories through integrating scattered contextual information from multiple interconnected memories. Given a question, OmniQuery retrieves relevant augmented memories and uses a large language model (LLM) to generate answers with references. In human evaluations, we show the effectiveness of OmniQuery with an accuracy of 71.5%, outperforming a conventional RAG system by winning or tying for 74.5% of the time.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47468b4b-1c77-445d-9d86-ad3d348793bbCited by top-tier papers4
- Creating General User Models from Computer UseOmar Shaikh, Shardul Sapkota, Shan Rizvi, Eric Horvitz et al.UIST 2025 · 13 citations
- ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable DevicesKevin Pu, Ting Zhang, Naveen Sendhilnathan, Sebastian Freitag et al.UIST 2025 · 9 citations
- StorySage: Conversational Autobiography Writing Powered by a Multi-Agent FrameworkShayan Talaei, Meijin Li, Kanu Grover, James Kent Hippler et al.UIST 2025 · 4 citations
- From Speech-to-Spatial: Grounding Utterances on A Live Shared View with Augmented RealityYoonsang Kim, Divyansh Pradhan, Devshree Jadeja, Arie E. KaufmanIEEE VR 2026
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-ReflectionAkari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil et al.ICLR 2024 · 1,798 citations
Related papers
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
- Crafting Personalized Agents through Retrieval-Augmented Generation on Editable Memory GraphsZheng Wang, Zhongyang Li, Zeren Jiang, Dandan Tu et al.EMNLP 2024 · 6 citations
- MemInsight: Autonomous Memory Augmentation for LLM AgentsRana Salama, Jason Cai, Michelle Yuan, Anna Currey et al.EMNLP 2025 · 1 citation
- SegMem-RAG: Adaptive Memory for Retrieval-Augmented Generation in Open-Ended Knowledge EnvironmentsXuanbo Fan, Tianqi Zhao, Yi Cheng, Chi Xiu et al.AAAI 2026
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language ModelsBernal Jimenez Gutierrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga et al.NeurIPS 2024 · 395 citations
