Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
Ziyi Bai, Ruiping Wang, Xilin Chen
Abstract
Video Question Answering (VideoQA) has emerged as a vital tool to evaluate agents' ability to understand human daily behaviors. Despite the recent success of large vision language models in many multi-modal tasks, complex situation reasoning over videos involving multiple human-object interaction events still remains challenging. In contrast, humans can easily tackle it by using a series of episode memories as anchors to quickly locate question-related key moments for reasoning. To mimic this effective reasoning strategy, we propose the Glance-Focus model. One simple way is to apply an action detection model to predict a set of actions as key memories. However, these actions within a closed set vocabulary are hard to generalize to various video domains. Instead of that, we train an Encoder-Decoder to generate a set of dynamic event memories at the glancing stage. Apart from using supervised bipartite matching to obtain the event memories, we further design an unsupervised memory generation method to get rid of dependence on event annotations. Next, at the focusing stage, these event memories act as a bridge to establish the correlation between the questions with high-level event concepts and low-level lengthy video content. Given the question, the model first focuses on the generated key event memory, then focuses on the most relevant moment for reasoning through our designed multi-level crossattention mechanism. We conduct extensive experiments on four Multi-Event VideoQA benchmarks including STAR, EgoTaskQA, AGQA, and NExT-QA. Our proposed model achieves state-of-the-art results, surpassing current large models in various challenging reasoning tasks. The code and models are available at https://github.com/ByZ0e/Glance-Focus . Recently, the emerged large vision-language models [1, 3, 8, 25, 47, 50] have demonstrated impressive generalization capabilities in VideoQA tasks. By leveraging pre-training on large-scale cross-modal pairs, these models can establish correlations between questions and videos to find answers effectively. Such approaches proves effective for short videos containing single or few events with one-step 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49ac60a0-c111-402f-a03b-83b0b5fb3b79Cited by top-tier papers4
- Slot-VLM: Object-Event Slots for Video-Language ModelingJiaqi Xu, Cuiling Lan, Wenxuan Xie, Xuejin Chen et al.NeurIPS 2024 · 13 citations
- Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning ScenariosShantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston TanNeurIPS 2024 · 9 citations
- Question-Answering Dense Video EventsHangyu Qin, Junbin Xiao, Angela YaoSIGIR 2025 · 5 citations
- Localizing Events in Videos with Multimodal QueriesGengyuan Zhang, Mang Ling Ada Fok, Jialu Ma, Yan Xia et al.CVPR 2025
Builds on26
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- SlowFast Networks for Video RecognitionChristoph Feichtenhofer, Haoqi Fan, Jitendra Malik, Kaiming HeICCV 2019 · 4,104 citations
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 1,624 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsAntoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi et al.ICCV 2019 · 1,437 citations
Related papers
- MoReVQA: Exploring Modular Reasoning Models for Video Question AnsweringJuhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho et al.CVPR 2024 · 27 citations
- GlFoMR: A Glance-then-Focus Multimodal Reasoning Framework for Diagram Question AnsweringYaxian Wang, Bifan Wei, Jun Liu, Lingling Zhang et al.SIGIR 2025
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented PerspectiveYan Zhang, Gangyan Zeng, Daiqing Wu, Huawen Shen et al.ACM MM 2025 · 2 citations
- MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringDifei Gao, Luowei Zhou, Lei Ji, Linchao Zhu et al.CVPR 2023
- Language-aware Visual Semantic Distillation for Video Question AnsweringBo Zou, Chao Yang, Yu Qiao, Chengbin Quan et al.CVPR 2024 · 2 citations
