Reasoning Models Are Test Exploiters: Rethinking Multiple Choice
Narun Raman, Taylor Lundy, Kevin Leyton-Brown
Abstract
When evaluating Large Language Models (LLMs) in question answering domains, it is common to ask the model to choose among a fixed set of choices (so-called multiplechoice question-answering, or MCQA). Although downstream tasks of interest typically do not provide systems with explicit options among which to choose, this approach is nevertheless widely used because it makes automatic grading straightforward and has tended to produce challenging benchmarks that correlate sufficiently well with downstream performance. This paper investigates the extent to which this trend continues to hold for state-of-the-art reasoning models, describing a systematic evaluation of 15 different question-answering benchmarks (e.g., MMLU, GSM8K, MATH, STEER-ME) and 27 different LLMs (including small models such as Qwen-2.5 7B Instruct, mid-sized models such as Llama-3.3 70B Instruct, and large state-of-the-art models such as OpenAI's o3). For each model-benchmark pair, we considered 5 ways of presenting the model with questions, including variations on whether multiple choices were offered to the model at all; whether "none of the above" sometimes replaced the right answer; and whether the model was permitted to perform chain-of-thought reasoning before and/or after the choices were presented. MCQA remained a good proxy for the downstream performance of models as long as they were allowed to perform chain-of-thought reasoning only before being presented with the options among which they had to select. On the other hand, large models that were able to perform reasoning after being given a set of options tended to significantly outperform their free-text performance due to exploiting the information in the options. We identify and quantify the signals models are using when answering MCQA questions, and offer practical guidelines when analyzing results from MCQA that better reflect LLMs' genuine reasoning capabilities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ABCD: All Biases Come DisguisedMateusz Nowak, Xavier Cadet, Peter ChinICML 2026 · 2 citations
- Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFTYesheng Liu, Hao Li, Haiyu Xu, Baoqi Pei et al.CVPR 2026 · 1 citation
Builds on5
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Large Language Models Are Not Robust Multiple Choice SelectorsChujie Zheng, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2024 · 424 citations
- STEER: Assessing the Economic Rationality of Large Language ModelsNarun Krishnamurthi Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine et al.ICML 2024 · 24 citations
- The Shifted and The Overlooked: A Task-oriented Investigation of User-GPT InteractionsSiru Ouyang, Shuohang Wang, Yang Liu, Ming Zhong et al.EMNLP 2023 · 10 citations
- Spurious Rewards: Rethinking Training Signals in RLVRRulin Shao, Stella Li, Rui Xin, Scott Geng et al.ICML 2026
Related papers
- Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?Nishant Balepur, Abhilasha Ravichander, Rachel RudingerACL 2024 · 5 citations
- Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?Neeladri Bhuiya, Viktor Schlegel, Stefan WinklerEMNLP 2024 · 2 citations
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
- How Long Reasoning Chains Influence LLMs' Judgment of Answer FactualityMinzhu Tu, Shiyu Ni, Keping BiACL 2026 · 5 citations
