What Evidence Do Language Models Find Convincing?
Alexander Wan, Eric Wallace, Dan Klein
Abstract
Retrieval-augmented language models are being increasingly tasked with subjective, contentious, and conflicting queries such as "is aspartame linked to cancer". To resolve these ambiguous queries, one must search through a large range of websites and consider which, if any, of this evidence do I find convincing? In this work, we study how LLMs answer this question. In particular, we construct CON-FLICTINGQA, a dataset that pairs controversial queries with a series of real-world evidence documents that contain different facts (e.g., quantitative results), argument styles (e.g., appeals to authority), and answers (Yes or No). We use this dataset to perform sensitivity and counterfactual analyses to explore which text features most affect LLM predictions. Overall, we find that current models rely heavily on the relevance of a website to the query, while largely ignoring stylistic features that humans find important such as whether a text contains scientific references or is written with a neutral tone. Taken together, these results highlight the importance of RAG corpus quality (e.g., the need to filter misinformation), and possibly even a shift in how LLMs are trained to better align with human judgements.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- SealQA: Raising the Bar for Reasoning in Search-Augmented Language ModelsThinh Pham, Nguyen Phan Nguyen, Pratibha Zunjare, Weiyuan Chen et al.ICLR 2026 · 69 citations
- Can Large Language Models Match the Conclusions of Systematic Reviews?Christopher Polzak, Alejandro Lozano, Min Woo Sun, James Burgess et al.ICLR 2026 · 9 citations
- NeuroGenPoisoning: Neuron-Guided Attacks on Retrieval-Augmented Generation of LLM via Genetic Optimization of External KnowledgeHanyu Zhu, Lance Fiondella, Jiawei Yuan, Kai Zeng et al.NeurIPS 2025 · 8 citations
- CUB: Benchmarking Context Utilisation Techniques for Language ModelsLovisa Hagström, Youna Kim, Haeun Yu, Sang-goo Lee et al.ACL 2026 · 5 citations
- Not All Contexts Are Equal: Teaching LLMs Credibility-aware GenerationRuotong Pan, Boxi Cao, Hongyu Lin, Xianpei Han et al.EMNLP 2024 · 4 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Efficiently Teaching an Effective Dense Retriever with Balanced Topic Aware SamplingSebastian Hofstätter, Sheng-Chieh Lin, Jheng-Hong Yang, Jimmy Lin et al.SIGIR 2021 · 297 citations
Related papers
- LLMs Trust Humans More, That's a Problem! Unveiling and Mitigating the Authority Bias in Retrieval-Augmented GenerationYuxuan Li, Xinwei Guo, Jiashi Gao, Guanhua Chen et al.ACL 2025
- Benchmarking LLM's Capability in Reasoning over Conflicting Web ReferencesYizhen Yuan, Rui Kong, Dongze Li, Yuanchun Li et al.ACL 2026
- SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QASher Badshah, Ali Emami, Hassan SajjadACL 2026 · 1 citation
- Conflict-Aware RAG: Multi-Stage Learning with Conflict Signals for Robust Retrieval-Augmented GenerationHaiyan Wu, Chenchen Wang, Chaoqun Sun, Chengxiong Lu et al.WWW 2026
- Search Arena: Analyzing Search-Augmented LLMsMihran Miroyan, Tsung-Han Wu, Logan King, Tianle Li et al.ICLR 2026 · 32 citations
