Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
Yuanchen Bei, Tianxin Wei, Xuying Ning, Yanjun Zhao, Zhining Liu, Xiao Lin, Yada Zhu, Hendrik F. Hamann, Jingrui He, Hanghang Tong
Abstract
Long-term memory is a critical capability for multimodal large language model (MLLM) agents, particularly in conversational settings where information accumulates and evolves over time. However, existing benchmarks either evaluate multi-session memory in textonly conversations or assess multimodal understanding within localized contexts, failing to evaluate how multimodal memory is preserved, organized, and evolved across longterm conversational trajectories. Thus, we introduce Mem-Gallery 1 , a new benchmark for evaluating multimodal long-term conversational memory in MLLM agents. Mem-Gallery features high-quality multi-session conversations grounded in both visual and textual information, with long interaction horizons and rich multimodal dependencies. Building on this dataset, we propose a systematic evaluation framework that assesses key memory capabilities along three functional dimensions: memory extraction and test-time adaptation, memory reasoning, and memory knowledge management. Extensive benchmarking across thirteen memory systems reveals several key findings, highlighting the necessity of explicit multimodal information retention and memory organization, the persistent limitations in memory reasoning and knowledge management, as well as the efficiency bottleneck of current models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 190a7091-596b-4708-9907-4d52f0d284bbCited by top-tier papers3
- Continual Low-Rank Adapters for LLM-based Generative Recommender SystemsHyunsik Yoo, Ting-Wei Li, SeongKu Kang, Zhining Liu et al.ICLR 2026 · 9 citations
- Prune as You Generate: Online Rollout Pruning for Faster and Better RLVRHaobo Xu, Sirui Chen, Ruizhong Qiu, Yuchen Yan et al.ACL 2026 · 6 citations
- Copyright-Bench: Agentic Evaluation of Copyright Law ComplianceZheng Hui, Doni Bloomfield, Noam KoltICML 2026 · 1 citation
Builds on5
- Evaluating Memory in LLM Agents via Incremental Multi-Turn InteractionsYuanzhe Hu, Yu Wang, Julian McAuleyICLR 2026 · 246 citations
- MuRAG: Multimodal Retrieval-Augmented Generator for Open Question Answering over Images and TextWenhu Chen, Hexiang Hu, Xi Chen, Pat Verga et al.EMNLP 2022 · 89 citations
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judgeDawei Li, Bohan Jiang, Liangjie Huang, Alimohammad Beigi et al.EMNLP 2025 · 37 citations
- Memory OS of AI AgentJiazheng Kang, Mingming Ji, Zhe Zhao, Ting BaiEMNLP 2025 · 4 citations
- Less is More: Empowering GUI Agent with Context-Aware SimplificationGongwei Chen, Xurui Zhou, Rui Shao, Yibo Lyu et al.ICCV 2025 · 2 citations
Related papers
- MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World ConversationHaochen Xue, Feilong Tang, Ming Hu, Yexin Liu et al.ACL 2025 · 23 citations
- ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional SupportTiantian Chen, Jiaqi Lu, Ying Shen, Lin ZhangWWW 2026 · 1 citation
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryLin Long, Yichen He, Wentao Ye, Yiyuan Pan et al.ICLR 2026 · 90 citations
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
- 3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language ModelWenbo Hu, Yining Hong, Yanjun Wang, Leison Gao et al.NeurIPS 2025 · 30 citations
