Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language Models
Md. Atabuzzaman, Ali Asgarov, Christopher Thomas
Abstract
Large Vision-Language Models (LVLMs) have achieved strong performance on visionlanguage tasks, particularly Visual Question Answering (VQA). While prior work has explored unimodal biases in VQA, the problem of selection bias in Multiple-Choice Question Answering (MCQA), where models may favor specific option tokens (e.g., "A") or positions, remains underexplored. In this paper, we investigate both the presence and nature of selection bias in LVLMs through fine-grained MCQA benchmarks spanning easy, medium, and hard difficulty levels, defined by the semantic similarity of the options. We further propose an inference-time logit-level debiasing method that estimates an ensemble bias vector from general and contextual prompts and applies confidence-adaptive corrections to the model's output. Our method mitigates bias without retraining and is compatible with frozen LVLMs. Extensive experiments across several state-ofthe-art models reveal consistent selection biases that intensify with task difficulty, and show that our mitigation approach significantly reduces bias while improving accuracy in challenging settings. This work offers new insights into the limitations of LVLMs in MCQA and presents a practical approach to improve their robustness in fine-grained visual reasoning. Datasets and code are available at: https://github.com/ Atabuzzaman/Selection-Bias-of-LVLMs * The proposed bias mitigation method was fully created and implemented by this author. Yellow-headed Blackbird B. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… C. This is a Rolls-Royce Phantom Drophead Coupe Convertible 2012 car which has a distinctive… D. This image shows a hot and sour soup food item which has a rich broth, egg strands, tofu, and … A. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… B. This image shows a Brown Creeper bird with small, dark eyes, a slender, upturned beak, and… C. This image shows a Florida Jay bird which has a small, blue-gray body, a white chest and belly… D. This image shows a Tree Sparrow bird which has a brown and white body, a chestnut cap … A. This image shows a Rusty Blackbird bird which has rusty-brown plumage in non-breeding… B This image shows a Red-winged Blackbird bird which has a black body, distinctive red and… C. This image shows a Brewer Blackbird bird which has a sleek black body, iridescent feathers… D. This image shows a Yellow-headed Blackbird bird which has a black body, bright yellow head… A. This is a Spyker C8 Coupe 2009 car which has a distinctive propeller logo, sleek aerodynamic … Hard Easy Medium
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8c9bf4f-534b-4e92-9750-1032390c8a54Builds on15
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- MUTANT: A Training Paradigm for Out-of-Distribution Generalization in Visual Question AnsweringTejas Gokhale, Pratyay Banerjee, Chitta Baral, Yezhou YangEMNLP 2020 · 136 citations
- Are Gender-Neutral Queries Really Gender-Neutral? Mitigating Gender Bias in Image SearchJialu Wang, Yang Liu, Xin Eric WangEMNLP 2021 · 44 citations
Related papers
- Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language ModelsOlga Loginova, Oleksandr Bezrukov, Ravi Shekhar, Alexey KravetsACL 2025 · 8 citations
- Adaptive Logit Adjustment for Debiasing Multimodal Language ModelsHoin Jung, Junyi Chai, Xiaoqian WangICLR 2026
- Language Bias in LVLMs: From In-Depth Analysis to Simple and Effective MitigationYangneng Chen, Jing LiICML 2026
- A Unified Debiasing Approach for Vision-Language Models across Modalities and TasksHoin Jung, Taeuk Jang, Xiaoqian WangNeurIPS 2024 · 27 citations
- MAVias: Mitigate any Visual BiasIoannis Sarridis, Christos Koutlis, Symeon Papadopoulos, Christos DiouICCV 2025 · 1 citation
