BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks
Nishant Balepur, Bhavya Rajasekaran, Hyunjin Jane Oh, Michael Xie, Atrey Desai, Vipul Gupta, Steven James Moore, Eunsol Choi, Rachel Rudinger, Jordan Lee Boyd-Graber
Abstract
Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-inspired toolkit using LLM judges to flag three common MCQ flaws: 1) contamination-items appearing exactly online; 2) shortcuts-cues in the choices that enable guessing; and 3) writing errors-structural/grammatical issues based on a 19-rule education rubric. We validate BenchMarker with human annotations, then run the tool to audit 12 benchmarks, revealing: 1) flaws persist in MCQA benchmarks, especially automatically-made and crowdsourced data-we detect 47% of TruthfulQA appears online and 100% of HellaSwag violates multiple writing rules; 2) contaminated MCQs tend to inflate accuracy, while writing errors tend to lower it and change rankings beyond random; and 3) prior benchmark repairs address their targeted issues (i.e., lowering accuracy with LLMwritten distractors), but inadvertently add new flaws (i.e. implausible distractors, many correct answers). Overall, flaws in MCQs degrade NLP evaluation, but education research offers a path forward. We release BenchMarker to bridge the fields and improve MCQA benchmark design. 1 Multiple-Choice Question j MCQ j … MCQ j … MCQ j … The appropriate place… www.quizlet.com Which can you drink? Inferred Question Rule 7: Is grammar consistent? 19-Rule Rubric from Education 1 (no question match) 0 (matches question) 1 (violates rule) 0 (no rule violation) Guess w/out question 1 (exists online) 0 (not online) + Contamination ( §3.1) Shortcuts ( §3.2) Writing Errors ( §3.3) Web search Report Card Contamination: 0.21 > Item j was found online Shortcuts: 0.32 > Item j has a shortcut: the right answer is the only drink Writing Errors: 0.45 > Item j R7: "This item" has a grammar mismatch -"plates" MCQA Benchmark BenchMarker: Highlighting Flaws Question: The appropriate place to put this item is a recycling bin Choices: (A) used motor oil (B) used soda can (C) used styrofoam plates (D) left over medicine Answer: (B)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 36c70df6-e222-4067-8359-ccc7ca94996dBuilds on20
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- QASC: A Dataset for Question Answering via Sentence CompositionTushar Khot, Peter Clark, Michal Guerquin, Peter Jansen et al.AAAI 2020 · 387 citations
- CRUXEval: A Benchmark for Code Reasoning, Understanding and ExecutionAlex Gu, Baptiste Rozière, Hugh James Leather, Armando Solar-Lezama et al.ICML 2024 · 270 citations
Related papers
- Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the AboveNishant Balepur, Rachel Rudinger, Jordan Lee Boyd-GraberACL 2025
- Does Question Really Matter? The Attribution of Answer Bias in LLM EvaluationBoxi Cao, Ruotong Pan, Hongyu Lin, Xianpei Han et al.AAAI 2026
- MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkQihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui et al.ACL 2025 · 33 citations
- Test of Time: Rethinking Temporal Signal of Benchmark ContaminationTerry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter et al.ACL 2026 · 3 citations
- Beyond Problem Solving: UOJ-Bench for Evaluating Code Generation, Hacking, and Repair in Competitive ProgrammingTingqiang Xu, Hangrui Zhou, Tianle Cai, Alex Gu et al.ICML 2026 · 1 citation
