Conflicting Needles in a Haystack: How LLMs behave when faced with contradictory information
Murathan Kurfali, Robert Östling
Abstract
Large Language Models (LLMs) have demonstrated an impressive ability to retrieve and summarize complex information, but their reliability in conflicting contexts remains poorly understood. We introduce an adversarial extension of the Needle-in-a-Haystack framework in which three mutually exclusive "needles" are embedded within long documents. By systematically manipulating factors such as position, repetition, layout, and domain relevance, we evaluate how LLMs handle contradictions. We find that models almost always fail to signal uncertainty and instead confidently select a single answer, exhibiting strong and consistent biases toward repetition, recency, and particular surface forms. We further analyze whether these patterns persist across model families and sizes, and we evaluate both probability-based and generation-based retrieval. Our framework highlights critical limitations in the robustness of current LLMs-including commercial systems-to contradiction. These limitations reveal potential shortcomings in RAG systems' ability to handle noisy or manipulated inputs and exposes risks for deployment in high-stakes applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c5c5e13-9bb9-431d-9662-3968e82e1249Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen et al.ICLR 2021 · 881 citations
- Benchmarking Large Language Models in Retrieval-Augmented GenerationJiawei Chen, Hongyu Lin, Xianpei Han, Le SunAAAI 2024 · 531 citations
- Retrieval meets Long Context Large Language ModelsPeng Xu, Wei Ping, Xianchao Wu, Lawrence McAfee et al.ICLR 2024 · 131 citations
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 31 citations
- DisentQA: Disentangling Parametric and Contextual Knowledge with Counterfactual Question AnsweringElla Neeman, Roee Aharoni, Or Honovich, Leshem Choshen et al.ACL 2023 · 20 citations
Related papers
- Summary of a Haystack: A Challenge to Long-Context LLMs and RAG SystemsPhilippe Laban, Alexander R. Fabbri, Caiming Xiong, Chien-Sheng WuEMNLP 2024 · 19 citations
- ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-SearchZeyu Shen, Basileal Imana, Tong Wu, Chong Xiang et al.NeurIPS 2025 · 26 citations
- Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesYizhen Yuan, Rui Kong, Dongze Li, Yuanchun Li et al.ACL 2026
- Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented GenerationHua Ye, Siyuan Chen, Ziqi Zhong, Canran Xiao et al.AAAI 2026 · 1 citation
- Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsFei Wang, Xingchen Wan, Ruoxi Sun, Jiefeng Chen et al.ACL 2025 · 50 citations
