Large Language Models Are Not Robust Multiple Choice Selectors
Chujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou, Minlie Huang
摘要
Multiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs). This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent "selection bias", namely, they prefer to select specific option IDs as answers (like "Option A"). Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs' token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs. To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model's prior bias for option IDs from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples, and then applies the estimated prior to debias the remaining samples. We demonstrate that it achieves interpretable and transferable debiasing with high computational efficiency. We hope this work can draw broader research attention to the bias and robustness of modern LLMs. 1 INTRODUCTION Question: In an SR latch built from NOR gates, which condition is not allowed Options: A.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper87
- On Prompt-Driven Safeguarding for Large Language ModelsChujie Zheng, Fan Yin, Hao Zhou, Fandong Meng 等ICML 2024 · 被引用 116 次
- Aligner: Efficient Alignment by Learning to CorrectJiaming Ji, Boyuan Chen, Hantao Lou, Donghai Hong 等NeurIPS 2024 · 被引用 115 次
- Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?Hyeong Kyu Choi, Xiaojin Zhu, Sharon LiNeurIPS 2025 · 被引用 93 次
- A Closer Look at the Limitations of Instruction TuningSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. 等ICML 2024 · 被引用 90 次
- Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic CorpusTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaNeurIPS 2024 · 被引用 60 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
相关 Paper
- Does Question Really Matter? The Attribution of Answer Bias in LLM EvaluationBoxi Cao, Ruotong Pan, Hongyu Lin, Xianpei Han 等AAAI 2026
- Benchmarking and Mitigating MCQA Selection Bias of Large Vision-Language ModelsMd. Atabuzzaman, Ali Asgarov, Christopher ThomasEMNLP 2025
- Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language ModelsOlga Loginova, Oleksandr Bezrukov, Ravi Shekhar, Alexey KravetsACL 2025 · 被引用 8 次
- Embracing Positional Bias in Multiple-Choice Question Answering via Permutation Equivariant Neural NetworksChengyu Jiao, Siyin Huang, Yu ZhangAAAI 2026
- Mitigating Selection Bias with Node Pruning and Auxiliary OptionsHyeong Kyu Choi, Weijie Xu, Chi Xue, Stephanie Eckman 等ACL 2025
