Memory-QA: Answering Recall Questions Based on Multimodal Memories
Hongda Jiang, Xinyuan Zhang, Siddhant Garg, Rishab Arora, Shiunzu Kuo, Jiayang Xu, Aaron Colak, Xin Luna Dong
Abstract
We introduce Memory-QA, a novel real-world task that involves answering recall questions about visual content from previously stored multimodal memories. This task poses unique challenges, including the creation of taskoriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, PENSIEVE , integrating memory-specific augmentation, time-and location-aware multi-signal retrieval, and multimemory QA fine-tuning. We created a multimodal benchmark to illustrate various real challenges in this task, and show the superior performance of PENSIEVE over state-of-the-art solutions (up to 14% on QA accuracy).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu et al.AAAI 2022 · 517 citations
Related papers
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga et al.EMNLP 2022 · 89 citations
- OmniQuery: Contextually Augmenting Captured Multimodal Memories to Enable Personal Question AnsweringJiahao Nick Li, Zhuohao Jerry Zhang, Jiaju MaCHI 2025 · 22 citations
- Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question AnsweringChangin Choi, Wonseok Lee, Jungmin Ko, Wonjong RheeACL 2026 · 2 citations
- Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge MemoryZiniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang et al.CVPR 2023
- Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search EnginesXinwei Long, Zhiyuan Ma, Ermo Hua, Kaiyan Zhang et al.AAAI 2025 · 18 citations
