Self-supervised Pre-training and Contrastive Representation Learning for Multiple-choice Video QA
Seonhoon Kim, Seohyeong Jeong, Eunbyul Kim, Inho Kang, Nojun Kwak
Abstract
Video Question Answering (Video QA) requires fine-grained understanding of both video and language modalities to answer the given questions. In this paper, we propose novel training schemes for multiple-choice video question answering with a self-supervised pre-training stage and a supervised contrastive learning in the main stage as an auxiliary learning. In the self-supervised pre-training stage, we transform the original problem format of predicting the correct answer into the one that predicts the relevant question to provide a model with broader contextual inputs without any further dataset or annotation. For contrastive learning in the main stage, we add a masking noise to the input corresponding to the ground-truth answer, and consider the original input of the ground-truth answer as a positive sample, while treating the rest as negative samples. By mapping the positive sample closer to the masked input, we show that the model performance is improved. We further employ locally aligned attention to focus more effectively on the video frames that are particularly relevant to the given corresponding subtitle sentences. We evaluate our proposed model on highly competitive benchmark datasets related to multiple-choice video QA: TVQA, TVQA+, and DramaQA. Experimental results show that our model achieves state-of-the-art performance on all datasets. We also validate our approaches through further analyses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- MERLOT: Multimodal Neural Script Knowledge ModelsRowan Zellers, Ximing Lu, Jack Hessel, Youngjae Yu et al.NeurIPS 2021 · 463 citations
- Zero-Shot Video Question Answering via Frozen Bidirectional Language ModelsAntoine Yang, Antoine Miech, Josef Sivic, Ivan Laptev et al.NeurIPS 2022 · 305 citations
- Tem-adapter: Adapting Image-Text Pretraining for Video Question AnswerGuangyi Chen, Xiao Liu, Guangrun Wang, Kun Zhang et al.ICCV 2023 · 32 citations
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh et al.EMNLP 2023 · 30 citations
- Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language ModelsJean Park, Kuk Jin Jang, Basam Alasaly, Sriharsha Mopidevi et al.AAAI 2025 · 21 citations
Builds on6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- TVQA+: Spatio-Temporal Grounding for Video Question AnsweringJie Lei, Licheng Yu, Tamara L. Berg, Mohit BansalACL 2020 · 173 citations
- Dense-Caption Matching and Frame-Selection Gating for Temporal Localization in VideoQAHyounghun Kim, Zineng Tang, Mohit BansalACL 2020 · 31 citations
- Modality Shifting Attention Network for Multi-Modal Video Question AnsweringJunyeong Kim, Minuk Ma, Trung X. Pham, Kyungsu Kim et al.CVPR 2020
Related papers
- Language-Guided Visual Aggregation Network for Video Question AnsweringXiao Liang, Di Wang, Quan Wang, Bo Wan et al.ACM MM 2023 · 5 citations
- Structured Video-Language Modeling with Temporal Grouping and Spatial GroundingYuanhao Xiong, Long Zhao, Boqing Gong, Ming-Hsuan Yang et al.ICLR 2024
- Video-Context Aligned Transformer for Video Question AnsweringLinlin Zong, Jiahui Wan, Xianchao Zhang, Xinyue Liu et al.AAAI 2024 · 8 citations
- Dual Hierarchical Temporal Convolutional Network with QA-Aware Dynamic Normalization for Video Story Question AnsweringFei Liu, Jing Liu, Xinxin Zhu, Richang Hong et al.ACM MM 2020 · 6 citations
- Multi-Question Learning for Visual Question AnsweringChenyi Lei, Lei Wu, Dong Liu, Zhao Li et al.AAAI 2020 · 9 citations
