Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers
Shijie Geng, Peng Gao, Moitreya Chatterjee, Chiori Hori, Jonathan Le Roux, Yongfeng Zhang, Hongsheng Li, Anoop Cherian
Abstract
Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a challenging multi-modal representation learning and reasoning scenario, advancements into which could influence several human-machine interaction applications. To solve this task, we introduce a semantics-controlled multi-modal shuffled Transformer reasoning framework, consisting of a sequence of Transformer modules, each taking a modality as input and producing representations conditioned on the input question. Our proposed Transformer variant uses a shuffling scheme on their multi-head outputs, demonstrating better regularization. To encode fine-grained visual information, we present a novel dynamic scene graph representation learning pipeline that consists of an intra-frame reasoning layer producing spatio-semantic graph representations for every frame, and an inter-frame aggregation module capturing temporal cues. Our entire pipeline is trained end-to-end. We present experiments on the benchmark AVSD dataset, both on answer generation and selection tasks. Our results demonstrate state-of-the-art performances on all evaluation metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 398a232b-8ba8-4085-b73f-f5ce790cf351Cited by top-tier papers10
- Path Language Modeling over Knowledge Graphsfor Explainable RecommendationShijie Geng, Zuohui Fu, Juntao Tan, Yingqiang Ge et al.WWW 2022 · 91 citations
- Container: Context Aggregation NetworksPeng Gao, Jiasen Lu, Hongsheng Li, Roozbeh Mottaghi et al.NeurIPS 2021 · 86 citations
- (2.5+1)D Spatio-Temporal Scene Graphs for Video Question AnsweringAnoop Cherian, Chiori Hori, Tim K. Marks, Jonathan Le RouxAAAI 2022 · 48 citations
- Efficient Non-Sampling Knowledge Graph EmbeddingZelong Li, Jianchao Ji, Zuohui Fu, Yingqiang Ge et al.WWW 2021 · 41 citations
- Constructing Holistic Spatio-Temporal Scene Graph for Video Semantic Role LabelingYu Zhao, Hao Fei, Yixin Cao, Bobo Li et al.ACM MM 2023 · 31 citations
Builds on2
Related papers
- SGEITL: Scene Graph Enhanced Image-Text Learning for Visual Commonsense ReasoningZhecan Wang, Haoxuan You, Liunian Harold Li, Alireza Zareian et al.AAAI 2022 · 40 citations
- AV-Dialog: Spoken Dialogue Models with Audio-Visual InputTuochao Chen, Bandhav Veluri, Hongyu Gong, Shyamnath GollakotaACL 2026 · 1 citation
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- Spatial-Temporal Transformer for Dynamic Scene Graph GenerationYuren Cong, Wentong Liao, Hanno Ackermann, Bodo Rosenhahn et al.ICCV 2021 · 163 citations
- Open-World 3D Scene Graph Generation for Retrieval-Augmented ReasoningFei Yu, Quan Deng, Shengeng Tang, Yuehua Li et al.AAAI 2026 · 2 citations
