Increasing Probability Mass on Answer Choices Does Not Always Improve Accuracy
Sarah Wiegreffe, Matthew Finlayson, Oyvind Tafjord, Peter Clark, Ashish Sabharwal
摘要
When pretrained language models (LMs) are applied to discriminative tasks such as multiplechoice questions, they place probability mass on vocabulary tokens that aren't among the given answer choices. Spreading probability mass across multiple surface forms with identical meaning (such as "bath" and "bathtub") is thought to cause an underestimation of a model's true performance, referred to as the "surface form competition" (SFC) hypothesis. This has motivated the introduction of various probability normalization methods. However, many core questions remain unanswered. How do we measure SFC? Are there direct ways of reducing it, and does doing so improve task performance? We propose a mathematical formalism for SFC which allows us to quantify and bound its impact for the first time. We identify a simple method for reducing it-namely, increasing probability mass on the given answer choices by a) including them in the prompt and b) using in-context learning with even just one example. We show this method eliminates the impact of SFC in the majority of instances. Our experiments on three diverse datasets and six LMs reveal several additional surprising findings. For example, both normalization and prompting methods for reducing SFC can be ineffective or even detrimental to task performance for some LMs. We conclude with practical insights for effectively prompting LMs for multiple-choice tasks. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Efficient semantic uncertainty quantification in language models via diversity-steered samplingJi Won Park, Kyunghyun ChoNeurIPS 2025 · 被引用 3 次
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice QuestionsSarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi 等ICLR 2025
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel 等ACL 2022 · 被引用 1,494 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 被引用 887 次
相关 Paper
- Surface Form Competition: Why the Highest Probability Answer Isn't Always RightAri Holtzman, Peter West, Vered Shwartz, Yejin Choi 等EMNLP 2021 · 被引用 10 次
- Answer-level Calibration for Free-form Multiple Choice Question AnsweringSawan KumarACL 2022 · 被引用 25 次
- Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' MemoryXingjian Tao, Yiwei Wang, Yujun Cai, Zhicheng Yang 等ICLR 2026 · 被引用 1 次
- Choices Speak Louder than QuestionsGyeongje Cho, Yeonkyoung So, Jaejin LeeICLR 2026 · 被引用 2 次
- Prompting Language Models for Linguistic StructureTerra Blevins, Hila Gonen, Luke ZettlemoyerACL 2023 · 被引用 15 次
