Seemingly Plausible Distractors in Multi-Hop Reasoning: Are Large Language Models Attentive Readers?
Neeladri Bhuiya, Viktor Schlegel, Stefan Winkler
Abstract
State-of-the-art Large Language Models (LLMs) are accredited with an increasing number of different capabilities, ranging from reading comprehension over advanced mathematical and reasoning skills to possessing scientific knowledge. In this paper we focus on multi-hop reasoning-the ability to identify and integrate information from multiple textual sources. Given the concerns with the presence of simplifying cues in existing multi-hop reasoning benchmarks, which allow models to circumvent the reasoning requirement, we set out to investigate whether LLMs are prone to exploiting such simplifying cues. We find evidence that they indeed circumvent the requirement to perform multi-hop reasoning, but they do so in more subtle ways than what was reported about their fine-tuned pre-trained language model (PLM) predecessors. We propose a challenging multi-hop reasoning benchmark by generating seemingly plausible multi-hop reasoning chains that ultimately lead to incorrect answers. We evaluate multiple open and proprietary state-of-the-art LLMs and show that their multi-hop reasoning performance is affected, as indicated by up to 45% relative decrease in F1 score when presented with such seemingly plausible alternatives. We also find that-while LLMs tend to ignore misleading lexical cues-misleading reasoning paths indeed present a significant challenge. The code and data are made available at https: //github.com/zawedcvg/Are-Large- Language-Models-Attentive-Readers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e654ed0e-a7a6-4be7-974a-dab3a866c489Cited by top-tier papers3
- Enhancing Reasoning Capabilities of LLMs via Principled Synthetic Logic CorpusTerufumi Morishita, Gaku Morio, Atsuki Yamaguchi, Yasuhiro SogawaNeurIPS 2024 · 60 citations
- Evaluating the Inductive Abilities of Large Language Models: Why Chain-of-Thought Reasoning Sometimes Hurts More Than HelpsHaibo Jin, Peiyan Zhang, Man Luo, Haohan WangNeurIPS 2025 · 1 citation
- Confidence Elicitation: A New Attack Vector for Large Language ModelsBrian Formento, Chuan-Sheng Foo, See-Kiong NgICLR 2025
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Large Language Models Can Be Easily Distracted by Irrelevant ContextFreda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales et al.ICML 2023 · 970 citations
Related papers
- CofCA: A STEP-WISE Counterfactual Multi-hop QA benchmarkJian Wu, Linyi Yang, Zhen Wang, Manabu Okumura et al.ICLR 2025
- MMhops-R1: Multimodal Multi-hop ReasoningTao Zhang, Ziqi Zhang, Zongyang Ma, Yuxin Chen et al.AAAI 2026
- Selection-Inference: Exploiting Large Language Models for Interpretable Logical ReasoningAntonia Creswell, Murray Shanahan, Irina HigginsICLR 2023 · 110 citations
- MMQA: Evaluating LLMs with Multi-Table Multi-Hop Complex QuestionsJian Wu, Linyi Yang, Dongyuan Li, Yuliang Ji et al.ICLR 2025
- Reasoning Models Are Test Exploiters: Rethinking Multiple ChoiceNarun Raman, Taylor Lundy, Kevin Leyton-BrownICML 2026 · 10 citations
