Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models
Olga Loginova, Oleksandr Bezrukov, Ravi Shekhar, Alexey Kravets
摘要
Evaluating Video Language Models (VLMs) is a challenging task. Due to its transparency, Multiple-Choice Question Answering (MCQA) is widely used to measure the performance of these models through accuracy. However, existing MCQA benchmarks fail to capture the full reasoning capabilities of VLMs due to selection bias, when models disproportionately favor certain answer options based on positional patterns observed during training. In this work, we conduct a comprehensive empirical analysis of several VLM architectures across major datasets designed to assess complex video-focused reasoning. We identify where the bias is most pronounced and demonstrate to what extent model responses reflect genuine understanding of video content and related questions, as opposed to reliance on arbitrary patterns or superficial cues, such as answer position. By decomposing the MCQA task and adapting fairness bias metrics to VLMs, we introduce a post-processing calibration technique BOLD to balance this bias. Our results show that reducing selection bias improves not only debiasing metrics but also overall model performance, including Accuracy and F1 Mean score. Our method, by suppressing"blind guessing", offers a more cost- and time-effective approach to mitigating selection bias compared to existing techniques. This study represents the first focused investigation of selection bias in video-to-text LLM-powered models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake DetectionJian-Yu Jiang-Lin, Kang-Yang Huang, Ling Zou, Ling Lo 等CVPR 2026 · 被引用 5 次
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect TimesOlga Loginova, Sofía Ortega LoguinovaACL 2025
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Self-Chained Image-Language Model for Video Localization and Question AnsweringShoubin Yu, Jaemin Cho, Prateek Yadav, Mohit BansalNeurIPS 2023 · 被引用 281 次
- Can I Trust Your Answer? Visually Grounded Video Question AnsweringJunbin Xiao, Angela Yao, Yicong Li, Tat-Seng ChuaCVPR 2024 · 被引用 44 次
- Revisiting the Role of Language Priors in Vision-Language ModelsZhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang 等ICML 2024 · 被引用 44 次
- Understanding Unimodal Bias in Multimodal Deep Linear NetworksYedi Zhang, Peter E. Latham, Andrew M. SaxeICML 2024 · 被引用 20 次
相关 Paper
- Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language ModelsMd. Atabuzzaman, Ali Asgarov, Christopher ThomasEMNLP 2025
- Large Language Models are Temporal and Causal Reasoners for Video Question AnsweringDohwan Ko, Ji Soo Lee, Woo-Young Kang, Byungseok Roh 等EMNLP 2023 · 被引用 30 次
- Commonsense Video Question Answering through Video-Grounded Entailment Tree ReasoningHuabin Liu, Filip Ilievski, Cees G. M. SnoekCVPR 2025
- Does Question Really Matter? The Attribution of Answer Bias in LLM EvaluationBoxi Cao, Ruotong Pan, Hongyu Lin, Xianpei Han 等AAAI 2026
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
