KnowIT VQA: Answering Knowledge-Based Questions about Videos
Noa Garcia, Mayu Otani, Chenhui Chu, Yuta Nakashima
摘要
We propose a novel video understanding task by fusing knowledge-based and video question answering. First, we introduce KnowIT VQA, a video dataset with 24,282 human-generated question-answer pairs about a popular sitcom. The dataset combines visual, textual and temporal coherence reasoning together with knowledge-based questions, which need of the experience obtained from the viewing of the series to be answered. Second, we propose a video understanding model by combining the visual and textual video content with specific knowledge about the show. Our main findings are: (i) the incorporation of knowledge produces outstanding improvements for VQA in video, and (ii) the performance on KnowIT VQA still lags well behind human accuracy, indicating its usefulness for studying current video modelling limitations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- Just Ask: Learning to Answer Questions from Millions of Narrated VideosAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等ICCV 2021 · 被引用 345 次
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev 等NeurIPS 2022 · 被引用 305 次
- Pano-AVQA: Grounded Audio-Visual Question Answering on 360° VideosHeeseung Yun, Youngjae Yu, Wonsuk Yang, Kangil Lee 等ICCV 2021 · 被引用 124 次
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu 等NeurIPS 2022 · 被引用 119 次
- Boosting Visual Question Answering with Context-aware Knowledge AggregationGuohao Li, Xin Wang, Wenwu ZhuACM MM 2020 · 被引用 82 次
相关 Paper
- On the hidden treasure of dialog in video question answeringDeniz Engin, François Schnitzler, Ngoc Q. K. Duong, Yannis AvrithisICCV 2021 · 被引用 12 次
- FriendsQA: A New Large-Scale Deep Video Understanding Dataset with Fine-grained Topic Categorization for Story VideosZhengqian Wu, Ruizhe Li, Zijun Xu, Zhongyuan Wang 等AAAI 2025 · 被引用 2 次
- DramaQA: Character-Centered Video Story Understanding with Hierarchical QASeongho Choi, Kyoung-Woon On, Yu-Jung Heo, Ahjeong Seo 等AAAI 2021 · 被引用 64 次
- From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-AnsweringJiangtong Li, Li Niu, Liqing ZhangCVPR 2022 · 被引用 48 次
- IntentQA: Context-aware Video Intent ReasoningJiapeng Li, Ping Wei, Wenjuan Han, Lifeng FanICCV 2023 · 被引用 97 次
