On the hidden treasure of dialog in video question answering
Deniz Engin, François Schnitzler, Ngoc Q. K. Duong, Yannis Avrithis
Abstract
High-level understanding of stories in video such as movies and TV shows from raw data is extremely challenging. Modern video question answering (VideoQA) systems often use additional human-made sources like plot synopses, scripts, video descriptions or knowledge bases. In this work, we present a new approach to understand the whole story without such external sources. The secret lies in the dialog: unlike any prior work, we treat dialog as a noisy source to be converted into text description via dialog summarization, much like recent methods treat video. The input of each modality is encoded by transformers independently, and a simple fusion method combines all modalities, using soft temporal attention for localization over long inputs. Our model outperforms the state of the art on the KnowIT VQA dataset by a large margin, without using question-specific human annotation or human-made plot summaries. It even outperforms human evaluators who have never watched any whole episode before. Code is available at https://engindeniz.github.io/dialogsummary-videoqa
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97bf7e46-924e-4131-8397-398312ba9149Cited by top-tier papers5
- Revisiting the "Video" in Video-Language UnderstandingShyamal Buch, Cristóbal Eyzaguirre, Adrien Gaidon, Jiajun Wu et al.CVPR 2022 · 121 citations
- Video Question Answering: Datasets, Algorithms and ChallengesYaoyao Zhong, Wei Ji, Junbin Xiao, Yicong Li et al.EMNLP 2022 · 70 citations
- i-Code: An Integrative and Composable Multimodal Learning FrameworkZiyi Yang, Yuwei Fang, Chenguang Zhu, Reid Pryzant et al.AAAI 2023 · 53 citations
- Modal-specific Pseudo Query Generation for Video Corpus Moment RetrievalMinjoon Jung, Seongho Choi, Joochan Kim, Jin-Hwa Kim et al.EMNLP 2022 · 11 citations
- Inferential Knowledge-Enhanced Integrated Reasoning for Video Question AnsweringJianguo Mao, Wenbin Jiang, Hong Liu, Xiangdong Wang et al.AAAI 2023 · 1 citation
Builds on6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
- Multi-View Sequence-to-Sequence Models with Conversational Structure for Abstractive Dialogue SummarizationJiaao Chen, Diyi YangEMNLP 2020 · 121 citations
- KnowIT VQA: Answering Knowledge-Based Questions about VideosNoa Garcia, Mayu Otani, Chenhui Chu, Yuta NakashimaAAAI 2020 · 93 citations
- Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQAHyounghun Kim, Zineng Tang, Mohit BansalACL 2020 · 31 citations
Related papers
- Vx2Text: End-to-End Learning of Video-Based Text Generation From Multimodal InputsXudong Lin, Gedas Bertasius, Jue Wang, Shih-Fu Chang et al.CVPR 2021
- CLIP-It! Language-Guided Video SummarizationMedhini Narasimhan, Anna Rohrbach, Trevor DarrellNeurIPS 2021 · 196 citations
- Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption GenerationKarina Abubakirova, Waseem Ullah, Latif U. Khan, Mohsen GuizaniKDD 2026
- VSTAR: A Video-grounded Dialogue Dataset for Situated Semantic Understanding with Scene and Topic TransitionsYuxuan Wang, Zilong Zheng, Xueliang Zhao, Jinpeng Li et al.ACL 2023 · 4 citations
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie et al.CVPR 2026 · 4 citations
