Open-Vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models
Dohwan Ko, Ji Soo Lee, Miso Choi, Jaewon Chu, Jihwan Park, Hyunwoo J. Kim
Abstract
Video Question Answering (VideoQA) is a challenging task that entails complex multi-modal reasoning. In contrast to multiple-choice VideoQA which aims to predict the answer given several options, the goal of open-ended VideoQA is to answer questions without restricting candidate answers. However, the majority of previous VideoQA models formulate open-ended VideoQA as a classification task to classify the video-question pairs into a fixed answer set, i.e., closed-vocabulary, which contains only frequent answers (e.g., top-1000 answers). This leads the model to be biased toward only frequent answers and fail to generalize on out-of-vocabulary answers. We hence propose a new benchmark, Open-vocabulary Video Question Answering (OVQA), to measure the generalizability of VideoQA models by considering rare and unseen answers. In addition, in order to improve the model’s generalization power, we introduce a novel GNN-based soft verbalizer that enhances the prediction on rare and unseen answers by aggregating the information from their similar words. For evaluation, we introduce new baselines by modifying the existing (closed-vocabulary) open-ended VideoQA models and improve their performances by further taking into account rare and unseen answers. Our ablation studies and qualitative analyses demonstrate that our GNN-based soft verbalizer further improves the model performance, especially on rare and unseen answers. We hope that our benchmark OVQA can serve as a guide for evaluating the generalizability of VideoQA models and inspire future research. Code is available at https://github.com/mlvlab/OVQA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8849abae-4ed6-4f30-82b3-533e2707236dCited by top-tier papers7
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh et al.EMNLP 2023 · 30 citations
- ViKey: Enhancing Temporal Understanding in Videos via Visual PromptingYeonkyung Lee, Dayun Ju, Youngmin Kim, Seil Kang et al.CVPR 2026 · 3 citations
- MoE-GRPO: Optimizing Mixture-of-Experts via Reinforcement Learning in Vision-Language ModelsDohwan Ko, Jinyoung Park, Seoung Choi, Sanghyeok Lee et al.CVPR 2026 · 3 citations
- TOGA: Temporally Grounded Open-Ended Video QA with Weak SupervisionAyush Gupta, Anirban Roy, Rama Chellappa, Nathaniel D. Bastian et al.ICCV 2025 · 2 citations
- Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-AnsweringZhaohe Liao, Jiangtong Li, Li Niu, Liqing ZhangCVPR 2024
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
Related papers
- When Open-Vocabulary Visual Question Answering Meets Causal Adapter: Benchmark and ApproachFeifei Zhang, Zhaoyi Zhang, Xi Zhang, Changsheng XuAAAI 2025
- NExT-QA: Next Phase of Question-Answering to Explaining Temporal ActionsJunbin Xiao, Xindi Shang, Angela Yao, Tat-Seng ChuaCVPR 2021
- ANetQA: A Large-scale Benchmark for Fine-grained Compositional Reasoning over Untrimmed VideosZhou Yu, Lixiang Zheng, Zhou Zhao, Fei Wu et al.CVPR 2023
- Knowledge-Constrained Answer Generation for Open-Ended Video Question AnsweringYao Jin, Guocheng Niu, Xinyan Xiao, Jian Zhang et al.AAAI 2023 · 12 citations
- MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering BenchmarkShaden Shaar, Bradon Thymes, Sirawut Chaixanien, Claire Cardie et al.CVPR 2026 · 4 citations
