Structured Co-reference Graph Attention for Video-grounded Dialogue
Junyeong Kim, Sunjae Yoon, Dahyun Kim, Chang D. Yoo
摘要
A video-grounded dialogue system referred to as the Structured Co-reference Graph Attention (SCGA) is presented for decoding the answer sequence to a question regarding a given video while keeping track of the dialogue context. Although recent efforts have made great strides in improving the quality of the response, performance is still far from satisfactory. The two main challenging issues are as follows: (1) how to deduce co-reference among multiple modalities and (2) how to reason on the rich underlying semantic structure of video with complex spatial and temporal dynamics. To this end, SCGA is based on (1) Structured Co-reference Resolver that performs dereferencing via building a structured graph over multiple modalities, (2) Spatio-temporal Video Reasoner that captures local-to-global dynamics of video via gradually neighboring graph attention. SCGA makes use of pointer network to dynamically replicate parts of the question for decoding the answer sequence. The validity of the proposed SCGA is demonstrated on AVSD@DSTC7 and AVSD@DSTC8 datasets, a challenging video-grounded dialogue benchmarks, and TVQA dataset, a large-scale videoQA benchmark. Our empirical results show that SCGA outperforms other state-of-the-art dialogue systems on both benchmarks, while extensive ablation study and qualitative analysis reveal performance gain and improved interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- MS-DETR: Natural Language Video Localization with Sampling Moment-Moment InteractionJing Wang, Aixin Sun, Hao Zhang, Xiaoli LiACL 2023 · 被引用 13 次
- Information-Theoretic Text Hallucination Reduction for Video-grounded DialogueSunjae Yoon, Eunseop Yoon, Hee Suk Yoon, Junyeong Kim 等EMNLP 2022 · 被引用 10 次
- Fast and Highly Expressive Policy Learning for Offline Reinforcement Learning via Bootstrapped Flow Q-LearningThanh Nguyen, Tri Ton, Hongbin Choe, Minh-Tung Luu 等ICML 2026 · 被引用 3 次
- PDCR: Perception-Decomposed Confidence Reward for Vision-Language ReasoningHee Suk Yoon, Eunseop Yoon, Ji Woo Hong, SooHwan Eom 等CVPR 2026 · 被引用 3 次
- Video-Based Optimal Transport for Feedback-Efficient Offline Preference-Based Reinforcement LearningMinh-Tung Luu, Hwanhee Kim, Younghwan Lee, Chang D. YooICML 2026 · 被引用 1 次
它引用的顶会 Paper4
- Relation-Aware Graph Attention Network for Visual Question AnsweringLinjie Li, Zhe Gan, Yu Cheng, Jingjing LiuICCV 2019 · 被引用 391 次
- Reasoning with Heterogeneous Graph Alignment for Video Question AnsweringPin Jiang, Yahong HanAAAI 2020 · 被引用 214 次
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim 等CVPR 2020
- Iterative Answer Prediction With Pointer-Augmented Multimodal Transformers for TextVQARonghang Hu, Amanpreet Singh, Trevor Darrell, Marcus RohrbachCVPR 2020
相关 Paper
- BiST: Bi-directional Spatio-Temporal Reasoning for Video-Grounded DialoguesHung Le, Doyen Sahoo, Nancy F. Chen, Steven C. H. HoiEMNLP 2020 · 被引用 30 次
- Dynamic Spatio-Temporal Modular Network for Video Question AnsweringZi Qian, Xin Wang, Xuguang Duan, Hong Chen 等ACM MM 2022 · 被引用 14 次
- Where Does It Exist: Spatio-Temporal Video Grounding for Multi-Form SentencesZhu Zhang, Zhou Zhao, Yang Zhao, Qi Wang 等CVPR 2020
- Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question AnsweringJianwen Jiang, Ziqiang Chen, Haojie Lin, Xibin Zhao 等AAAI 2020 · 被引用 129 次
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 被引用 173 次
