Dynamic Spatio-Temporal Modular Network for Video Question Answering
Zi Qian, Xin Wang, Xuguang Duan, Hong Chen, Wenwu Zhu
摘要
Video Question Answering (VideoQA) aims to understand given videos and questions comprehensively by generating correct answers. However, existing methods usually rely on end-to-end black-box deep neural networks to infer the answers, which significantly differs from human logic reasoning, thus lacking the ability to explain. Besides, the performances of existing methods tend to drop when answering compositional questions involving realistic scenarios. To tackle these challenges, we propose a Dynamic Spatio-Temporal Modular Network (DSTN) model, which utilizes a spatio-temporal modular network to simulate the compositional reasoning procedure of human beings. Concretely, we divide the task of answering a given question into a set of sub-tasks focusing on certain key concepts in questions and videos such as objects, actions, temporal orders, etc. Each sub-task can be solved with a separately designed module, e.g., spatial attention module, temporal attention module, logic module, and answer module. Then we dynamically assemble different modules assigned with different sub-tasks to generate a tree-structured spatio-temporal modular neural network for human-like reasoning before producing the final answer for the question. We carry out extensive experiments on the AGQA dataset to demonstrate our proposed DSTN model can significantly outperform several baseline methods in various settings. Moreover, we evaluate intermediate results and visualize each reasoning step to verify the rationality of different modules and the explainability of the proposed DSTN model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 被引用 97 次
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li 等EMNLP 2022 · 被引用 70 次
- STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question AnsweringYueqian Wang, Yuxuan Wang, Kai Chen, Dongyan ZhaoAAAI 2024 · 被引用 4 次
- Divide and Conquer: Exploring Language-centric Tree Reasoning for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Siyu Sun, Qingyang Liu 等ICML 2025
- Modularized Self-Reflected Video Reasoner for Multimodal LLM with Application to Video Question AnsweringZihan Song, Xin Wang, Zi Qian, Hong Chen 等ICML 2025
它引用的顶会 Paper18
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- CLEVRER: Collision Events for Video Representation and ReasoningKexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli 等ICLR 2020 · 被引用 584 次
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu 等NeurIPS 2021 · 被引用 463 次
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 被引用 214 次
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du 等AAAI 2020 · 被引用 187 次
相关 Paper
- Progressive Spatio-temporal Perception for Audio-Visual Question AnsweringGuangyao Li, Wenxuan Hou, Di HuACM MM 2023 · 被引用 39 次
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 等CVPR 2020
- Dual Hierarchical Temporal Convolutional Network with QA-Aware Dynamic Normalization for Video Story Question AnsweringFei Liu, Jing Liu, Xinxin Zhu, Richang Hong 等ACM MM 2020 · 被引用 6 次
- MIST : Multi-modal Iterative Spatial-Temporal Transformer for Long-form Video Question AnsweringDifei Gao, Luowei Zhou, Lei Ji, Linchao Zhu 等CVPR 2023
- Structured Co-reference Graph Attention for Video-grounded DialogueJunyeong Kim, Sunjae Yoon, Dahyun Kim, Chang D. YooAAAI 2021 · 被引用 31 次
