MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
Juhong Min, Shyamal Buch, Arsha Nagrani, Minsu Cho, Cordelia Schmid
摘要
This paper addresses the task of video question answering (videoQA) via a decomposed multi-stage, modular rea-soning framework. Previous modular methods have shown promise with a single planning stage ungrounded in visual content. However, through a simple and effective base-line, we find that such systems can lead to brittle behavior in practice for challenging videoQA settings. Thus, unlike traditional single-stage planning methods, we propose a multi-stage system consisting of an event parser, a grounding stage, and a final reasoning stage in conjunction with an external memory. All stages are training-free, and performed using few-shot prompting of large models, creating interpretable intermediate outputs at each stage. By decomposing the underlying planning and task complexity, our method, MoReVQA, improves over prior work on stan-dard videoQA benchmarks (NExT-QA, iVQA, EgoSchema, ActivityNet-QA) with state-of-the-art results, and extensions to related tasks (grounded videoQA, paragraph captioning).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-TuningQi (Cheems) Wang, Yanrui Yu, Ye Yuan, Rui Mao 等NeurIPS 2025 · 被引用 103 次
- Temporal Chain of Thought: Long-Video Understanding by Thinking in FramesAnurag Arnab, Ahmet Iscen, Mathilde Caron, Alireza Fathi 等NeurIPS 2025 · 被引用 31 次
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data SynthesisCongzhi Zhang, Zhibin Wang, Yinchao Ma, Jiawei Peng 等ICLR 2026 · 被引用 24 次
- Time Blindness: Why Video-Language Models Can't See What Humans Can?Ujjwal Upadhyay, Mukul Ranjan, Zhiqiang Shen, Mohamed ElhoseinyCVPR 2026 · 被引用 17 次
- Threading Keyframe with Narratives: MLLMs as Strong Long Video ComprehendersBo Fang, Yuxin Song, Haoyuan Sun, Qiangqiang Wu 等ICLR 2026 · 被引用 13 次
它引用的顶会 Paper35
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
相关 Paper
- Question-Answering Dense Video EventsHangyu Qin, Junbin Xiao, Angela YaoSIGIR 2025 · 被引用 5 次
- M-LLM Based Video Frame Selection for Efficient Video UnderstandingKai Hu, Feng Gao, Xiaohan Nie, Peng Zhou 等CVPR 2025
- Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu 等ICML 2025
- Glance and Focus: Memory Prompting for Multi-Event Video Question AnsweringZiyi Bai, Ruiping Wang, Xilin ChenNeurIPS 2023 · 被引用 19 次
- A Simple LLM Framework for Long-Range Video Question-AnsweringCe Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang 等EMNLP 2024 · 被引用 37 次
