Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
Nishant Balepur, Abhilasha Ravichander, Rachel Rudinger
摘要
Multiple-choice question answering (MCQA) is often used to evaluate large language models (LLMs). To see if MCQA assesses LLMs as intended, we probe if LLMs can perform MCQA with choices-only prompts, where models must select the correct answer only from the choices. In three MCQA datasets and four LLMs, this prompt bests a majority baseline in 11/12 cases, with up to 0.33 accuracy gain. To help explain this behavior, we conduct an in-depth, black-box analysis on memorization, choice dynamics, and question inference. Our key findings are threefold. First, we find no evidence that the choices-only accuracy stems from memorization alone. Second, priors over individual choices do not fully explain choicesonly accuracy, hinting that LLMs use the group dynamics of choices. Third, LLMs have some ability to infer a relevant question from choices, and surprisingly can sometimes even match the original question. Inferring the original question is an impressive reasoning strategy, but it cannot fully explain the high choices-only accuracy of LLMs in MCQA. Thus, while LLMs are not fully incapable of reasoning in MCQA, we still advocate for the use of stronger baselines in MCQA benchmarks, the design of robust MCQA datasets for fair evaluations, and further efforts to explain LLM decision-making. 1 Question: Which of these contains only a solution? Answer: (B) Question: Which can be considered a solution? Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Question: Which can be considered a solution? Step 2: Answer the Question from Step 1 Classify Choice (A) Correctness Classify Choice (B) Correctness ... Abductive Question Inference ( §6) Question: Which of these contains only a solution? Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Choices: (A) a jar of pickles (B) a bottle of juice (C) a bag of peanuts (D) a can of mixed fruit Answer: (B) Full MCQA Prompt Choices-only Prompt ( §3) LLMs Can Perform MCQA with no Question, but how? Classify Choice (D) Correctness Question: Which of these contains only a solution? Choices: (A) (B) (C) (D) Answer: (B)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMsKyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak 等ICLR 2026 · 被引用 59 次
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM HallucinationsBuyun Liang, Liangzu Peng, Jinqi Luo, Darshan Thaker 等NeurIPS 2025 · 被引用 11 次
- Detecting Data Contamination in LLMs via In-Context LearningMichal Zawalski, Meriem Boubdir, Klaudia Balazy, Besmira Nushi 等ICLR 2026 · 被引用 8 次
- Susu Box or Piggy Bank: Assessing Cultural Commonsense Knowledge between Ghana and the USChristabel Acquaye, Haozhe An, Rachel RudingerEMNLP 2024 · 被引用 2 次
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal InconsistenciesLukas Selch, Yufang Hou, Muhammad Jehanzeb Mirza, Sivan Doveh 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper11
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
相关 Paper
- Reasoning Models Are Test Exploiters: Rethinking Multiple ChoiceNarun Raman, Taylor Lundy, Kevin Leyton-BrownICML 2026 · 被引用 10 次
- Leveraging Large Language Models for Multiple Choice Question AnsweringJoshua Robinson, David WingateICLR 2023 · 被引用 40 次
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Does Question Really Matter? The Attribution of Answer Bias in LLM EvaluationBoxi Cao, Ruotong Pan, Hongyu Lin, Xianpei Han 等AAAI 2026
- Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language ModelsOlga Loginova, Oleksandr Bezrukov, Ravi Shekhar, Alexey KravetsACL 2025 · 被引用 8 次
