BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded Dialogues
Hung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. Hoi
Abstract
Video-grounded dialogues are very challenging due to (i) the complexity of videos which contain both spatial and temporal variations, and (ii) the complexity of user utterances which query different segments and/or different objects in videos over multiple dialogue turns. However, existing approaches to video-grounded dialogues often focus on superficial temporal-level visual cues, but neglect more fine-grained spatial signals from videos. To address this drawback, we propose Bi-directional Spatio-Temporal Learning (BiST), a vision-language neural framework for high-resolution queries in videos based on textual cues. Specifically, our approach not only exploits both spatial and temporal-level information, but also learns dynamic information diffusion between the two feature spaces through spatial-to-temporal and temporal-tospatial reasoning. The bidirectional strategy aims to tackle the evolving semantics of user queries in the dialogue setting. The retrieved visual cues are used as contextual information to construct relevant responses to the users. Our empirical results and comprehensive qualitative analysis show that BiST achieves competitive performance and generates reasonable responses on a large-scale AVSD benchmark. We also adapt our BiST models to the Video QA setting, and substantially outperform prior approaches on the TGIF-QA benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Learning to Retrieve Videos by Asking QuestionsAvinash Madasu, Junier Oliva, Gedas BertasiusACM MM 2022 · 17 citations
- META-GUI: Towards Multi-modal Conversational Agents on Mobile GUILiangtai Sun, Xingyu Chen, Lu Chen, Tianle Dai et al.EMNLP 2022 · 11 citations
- FashionNTM: Multi-turn Fashion Image Retrieval via Cascaded MemoryAnwesan Pal, Sahil Wadhwa, Ayush Jaiswal, Xu Zhang et al.ICCV 2023 · 11 citations
- DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded DialogueHung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami et al.ACL 2021
Builds on5
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 214 citations
- Location-Aware Graph Convolutional Networks for Video Question AnsweringDeng Huang, Peihao Chen, Runhao Zeng, Qing Du et al.AAAI 2020 · 187 citations
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao et al.AAAI 2020 · 129 citations
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li et al.AAAI 2020 · 9 citations
- Hierarchical Conditional Relation Networks for Video Question AnsweringThao Minh Le, Vuong Le, Svetha Venkatesh, Truyen TranCVPR 2020
Related papers
- Structured Co-reference Graph Attention for Video-grounded DialogueJunyeong Kim, Sunjae Yoon, Dahyun Kim, Chang D. YooAAAI 2021 · 31 citations
- Collaborative Static and Dynamic Vision-Language Streams for Spatio-Temporal Video GroundingZihang Lin, Chaolei Tan, Jian-Fang Hu, Zhi Jin et al.CVPR 2023
- Learning Reasoning Paths over Semantic Graphs for Video-grounded DialoguesHung Le, Nancy F. Chen, Steven C. H. HoiICLR 2021 · 18 citations
- VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal SegmentationJihwan Hong, Jaeyoung DoCVPR 2026 · 2 citations
- Agentic Spatio-Temporal Grounding via Collaborative ReasoningHeng Zhao, Yew-Soon Ong, Joey Tianyi ZhouSIGIR 2026 · 1 citation
