Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
Md. Atabuzzaman, Ali Asgarov, Christopher Thomas
摘要
Large Vision-Language Models (LVLMs) have achieved strong performance on visionlanguage tasks, particularly Visual Question Answering (VQA). While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor specific option tokens (e.g., "A") or positions, remains underexplored. In this paper, we investigate both the presence and nature of selection bias in LVLMs through fine-grained MCQA benchmarks spanning easy, medium, and hard difficulty levels, defined by the semantic similarity of the options. We further propose an inference-time logit-level debiasing method that estimates an ensemble bias vector from general and contextual prompts and applies confidence-adaptive corrections to the model's output. Our method mitigates bias without retraining and is compatible with frozen LVLMs. Extensive experiments across several state-ofthe-art models reveal consistent selection biases that intensify with task difficulty, and show that our mitigation approach significantly reduces bias while improving accuracy in challenging settings. This work offers new insights into the limitations of LVLMs in MCQA and presents a practical approach to improve their robustness in fine-grained visual reasoning. Datasets and code are available at: https://github.com/ Atabuzzaman/Selection-Bias-of-LVLMs * The proposed bias mitigation method was fully created and implemented by this author. Yellow-headed Blackbird B. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… C. This is a Rolls-Royce Phantom Drophead Coupe Convertible 2012 car which has a distinctive… D. This image shows a hot and sour soup food item which has a rich broth, egg strands, tofu, and … A. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… B. This image shows a Brown Creeper bird with small, dark eyes, a slender, upturned beak, and… C. This image shows a Florida Jay bird which has a small, blue-gray body, a white chest and belly… D. This image shows a Tree Sparrow bird which has a brown and white body, a chestnut cap … A. This image shows a Rusty Blackbird bird which has rusty-brown plumage in non-breeding… B This image shows a Red-winged Blackbird bird which has a black body, distinctive red and… C. This image shows a Brewer Blackbird bird which has a sleek black body, iridescent feathers… D. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… A. This is a Spyker C8 Coupe 2009 car which has a distinctive propeller logo, sleek aerodynamic … Hard Easy Medium
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou 等ICLR 2024 · 被引用 424 次
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 被引用 136 次
- Are Gender-Neutral Queries Really Gender-Neutral? Mitigating Gender Bias in Image SearchJialu Wang, Yang Liu, Xin Eric WangEMNLP 2021 · 被引用 44 次
相关 Paper
- Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language ModelsOlga Loginova, Oleksandr Bezrukov, Ravi Shekhar, Alexey KravetsACL 2025 · 被引用 8 次
- Adaptive Logit Adjustment for Debiasing Multimodal Language ModelsHoin Jung, Junyi Chai, Xiaoqian WangICLR 2026
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- A Unified Debiasing Approach for Vision-Language Models across Modalities and TasksHoin Jung, Taeuk Jang, Xiaoqian WangNeurIPS 2024 · 被引用 27 次
- MAVias: Mitigate any Visual BiasIoannis Sarridis, Christos Koutlis, Symeon Papadopoulos, Christos DiouICCV 2025 · 被引用 1 次
