Learning Reasoning Paths over Semantic Graphs for Video-grounded Dialogues
Hung Le, Nancy F. Chen, Steven C. H. Hoi
Abstract
Compared to traditional visual question answering, video-grounded dialogues require additional reasoning over dialogue context to answer questions in a multi-turn setting. Previous approaches to video-grounded dialogues mostly use dialogue context as a simple text input without modelling the inherent information flows at the turn level. In this paper, we propose a novel framework of Reasoning Paths in Dialogue Context (PDC). PDC model discovers information flows among dialogue turns through a semantic graph constructed based on lexical components in each question and answer. PDC model then learns to predict reasoning paths over this semantic graph. Our path prediction model predicts a path from the current turn through past dialogue turns that contain additional visual cues to answer the current question. Our reasoning model sequentially processes both visual and textual information through this reasoning path and the propagated features are used to generate the answer. Our experimental results demonstrate the effectiveness of our method and provide additional insights on how models use semantic dependencies in a dialogue context to retrieve visual cues.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5ca50a9d-25c9-4464-8a86-2e44ee75e5b5Cited by top-tier papers1
Ask how each one uses itBuilds on4
- Learning to Retrieve Reasoning Paths over Wikipedia Graph for Question AnsweringAkari Asai, Kazuma Hashimoto, Hannaneh Hajishirzi, Richard Socher et al.ICLR 2020 · 322 citations
- Who Did They Respond to? Conversation Structure Modeling Using Masked Hierarchical TransformerHenghui Zhu, Feng Nan, Zhiguo Wang, Ramesh Nallapati et al.AAAI 2020 · 41 citations
- Learning to Ask More: Semi-Autoregressive Sequential Question Generation under Dual-Graph InteractionZi Chai, Xiaojun WanACL 2020 · 26 citations
- Detecting Attackable Sentences in ArgumentsYohan Jo, Seojin Bang, Emaad A. Manzoor, Eduard H. Hovy et al.EMNLP 2020 · 25 citations
Related papers
- DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual DialogFeilong Chen, Fandong Meng, Jiaming Xu, Peng Li et al.AAAI 2020 · 35 citations
- DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual DialogueXiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang et al.AAAI 2020 · 72 citations
- DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded DialogueHung Le, Chinnadhurai Sankar, Seungwhan Moon, Ahmad Beirami et al.ACL 2021
- BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded DialoguesHung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. HoiEMNLP 2020 · 30 citations
- Inferential Knowledge-Enhanced Integrated Reasoning for Video Question AnsweringJianguo Mao, Wenbin Jiang, Hong Liu, Xiangdong Wang et al.AAAI 2023 · 1 citation
