KnowIT VQA: Answering Knowledge-Based Questions about Videos
Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima
Abstract
We propose a novel video understanding task by fusing knowledge-based and video question answering. First, we introduce KnowIT VQA, a video dataset with 24,282 human-generated question-answer pairs about a popular sitcom. The dataset combines visual, textual and temporal coherence reasoning together with knowledge-based questions, which need of the experience obtained from the viewing of the series to be answered. Second, we propose a video understanding model by combining the visual and textual video content with specific knowledge about the show. Our main findings are: (i) the incorporation of knowledge produces outstanding improvements for VQA in video, and (ii) the performance on KnowIT VQA still lags well behind human accuracy, indicating its usefulness for studying current video modelling limitations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.ICCV 2021 · 345 citations
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360° VideosHeeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee et al.ICCV 2021 · 124 citations
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu et al.NeurIPS 2022 · 119 citations
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 82 citations
Related papers
- On the hidden treasure of dialog in video question answeringDeniz Engin, François Schnitzler, Ngoc Q. K. Duong, Yannis AvrithisICCV 2021 · 12 citations
- FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story VideosZhengqian Wu, Ruizhe Li, Zijun Xu, Zhongyuan Wang et al.AAAI 2025 · 2 citations
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QASeongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo et al.AAAI 2021 · 64 citations
- From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringJiangtong Li, Li Niu, Liqing ZhangCVPR 2022 · 48 citations
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 97 citations
