DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question Answering
Borui Li, Xingcai Zhang, Tianen Liu, Shuai Wang, Yun Cheng, Shuai Wang
Abstract
Video question answering (VideoQA), whose goal is to produce answers through the integration of linguistic and visual understanding, has emerged as a significant research focus. Although Large Multimodal Models (LMMs) and autonomous agent methods have achieved notable advances in VideoQA, excessive computational overhead and restricted multimodal interaction capabilities limit their ability to facilitate the continuous evolution of the VideoQA system. To address the challenge, we introduce DigimonGPT, an evolvable VideoQA agent inspired by cognitive psychology. Specifically, DigimonGPT integrates a multimodal memory mechanism to achieve the continuous evolution of VideoQA systems. An intra-video declarative memory contains fundamental features of the video and semantic contexts extracted from historical QA pairs. Another inter-task procedural memory encodes task-solving experience for further question answering. Additionally, we introduce a hierarchical memory replay mechanism for VideoQA that selects appropriate memories by their relevance and question complexity. Extensive experiments demonstrate that DigimonGPT's accuracy outperforms by an average of 13.71% on NExT-QA datasets and 9.89% on Intent-QA datasets over LMM and autonomous agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7e9bac3-f8ad-4a35-b4ce-8fa78a5f5be6Builds on30
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
- ViperGPT: Visual Inference via Python Execution for ReasoningDídac Surís, Sachit Menon, Carl VondrickICCV 2023 · 732 citations
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li et al.ICLR 2024 · 467 citations
Related papers
- Bridging the Grounding Gap in VideoQA via Typed Memory for Language-based Belief-State ReasoningSaman Forouzandeh, Wei Peng, Xinghuo Yu, Mahdi JaliliICML 2026
- MoReVQA: Exploring Modular Reasoning Models for Video Question AnsweringJuhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho et al.CVPR 2024 · 27 citations
- DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)Zongxin Yang, Guikun Chen, Xiaodi Li, Wenguan Wang et al.ICML 2024 · 70 citations
- WorldMM: Dynamic Multimodal Memory Agent for Long Video ReasoningWoongyeong Yeo, Kangsan Kim, Jaehong Yoon, Sung Ju HwangCVPR 2026 · 53 citations
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term MemoryLin Long, Yichen He, Wentao Ye, Yiyuan Pan et al.ICLR 2026 · 90 citations
